Method, device, electronic equipment and computer program product for processing financial transactions

CN122531385APending Publication Date: 2026-08-07中国工商银行股份有限公司山西省分行
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中国工商银行股份有限公司山西省分行
Filing Date
2026-04-07
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]本申请的主要目的在于提供一种金融业务的处理方法、装置、电子设备及计算机程序产品,以解决相关技术中通过语音指令的方式处理金融业务时存在语音识别精确度低、业务处理效率低的技术问题

Benefits of technology

[0018] In this embodiment, a financial transaction processing method is adopted. This involves receiving voice commands sent by a target user through a client, where the voice commands instruct the target user on the financial transaction they wish to apply for. The voice commands are then input into a speech recognition model to obtain speech-recognized text. The speech recognition model converts the voice commands into speech-recognized text. The type of financial transaction is determined based on the speech-recognized text. A target business model is then invoked based on the financial transaction type, and the target business model is controlled to execute the financial transaction associated with the voice command. This method solves the technical problems of low speech recognition accuracy and low business processing efficiency in related technologies when processing financial transactions via voice commands. By inputting voice commands into a speech recognition model, obtaining speech-recognized text, invoking the target business model based on the financial transaction type, and controlling the target business model to execute the financial transaction associated with the voice command, the technical effect of improving the speech recognition accuracy and business processing efficiency of processing financial transactions via voice commands is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531385A_ABST
    Figure CN122531385A_ABST
Patent Text Reader

Abstract

The application discloses a kind of processing method, device, electronic equipment and computer program product of financial service.It relates to the field of financial technology, and the method comprises: receiving the voice instruction sent by target user through client, wherein the voice instruction indicates the financial service applied by target user;Voice instruction is input into speech recognition model, and voice recognition text is obtained by processing, wherein the speech recognition model is used to convert voice instruction to obtain voice recognition text;Determine the financial service type according to voice recognition text;According to the type of financial service, call target service model, and control target service model to execute the financial service associated with voice instruction.By the present application, the technical problems of low voice recognition accuracy and low business processing efficiency in related art when processing financial service through voice instruction are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial technology, and more specifically, to a method, apparatus, electronic device, and computer program product for processing financial transactions. Background Technology

[0002] As financial institutions advance their digital transformation, the complex counter services and back-office system projects become the core intersection of front-office services and middle and back-office support for financial institutions. The standardization and traceability of their process execution directly affect customer experience.

[0003] However, currently, financial institutions still rely on traditional, experience-driven models in these two key business scenarios, lacking systematic process management support tools. This has exposed a series of problems: First, both counter service processing and project implementation involve multi-system linkage, multiple transaction instructions, and serial operations, resulting in long process chains and fragmented steps. Different departments, branches, and employees have differing understandings of the operational standards for the same type of business, and the lack of a unified digital carrier for process standards leads to low efficiency in business progress. Second, in a high-intensity, multi-tasking work environment, employees need to handle multiple tasks such as customer communication simultaneously. Relying solely on personal memory and paper notes to manage complex operational processes easily leads to omissions of key steps, resulting in business rework, high resource consumption, and low knowledge reuse rates.

[0004] To address these issues, financial institutions have applied voice interaction models to the aforementioned service areas. However, in scenarios such as counter service and project deployment, they still face problems such as low accuracy in recognizing business keywords and susceptibility to noise interference from the counter environment. Furthermore, the model relies on cloud-based voice recognition services, which poses a risk of leakage of sensitive information and makes it difficult to directly integrate with process tools.

[0005] There is currently no effective solution to the technical problems of low voice recognition accuracy and low business processing efficiency when handling financial business through voice commands in related technologies. Summary of the Invention

[0006] The main objective of this application is to provide a method, apparatus, electronic device, and computer program product for processing financial transactions, in order to solve the technical problems of low voice recognition accuracy and low business processing efficiency when processing financial transactions through voice commands in related technologies.

[0007] To achieve the above objectives, according to one aspect of this application, a method for processing financial transactions is provided. The method includes: receiving a voice command sent by a target user through a client, wherein the voice command instructs the target user to apply for a financial transaction; inputting the voice command into a speech recognition model and processing it to obtain speech-recognized text, wherein the speech recognition model is used to convert the voice command to obtain speech-recognized text; determining the type of financial transaction based on the speech-recognized text; invoking a target business model based on the type of financial transaction, and controlling the target business model to execute the financial transaction associated with the voice command.

[0008] Optionally, before receiving the voice command sent by the target user through the client, the method further includes: detecting whether the target user generates a trigger command through the client, wherein the trigger command is generated when the target user triggers a preset wake word through the client, or when the target user is in a preset scenario; if the trigger command is detected to be generated by the target user through the client, the step of receiving the voice command sent by the target user through the client is performed.

[0009] Optionally, receiving voice commands sent by the target user through the client includes: receiving an initial voice command sent by the target user through the client with the authorization of the target user; performing silence detection on the initial voice command; if noise exists in the initial voice command, performing noise suppression on the initial voice command to obtain a suppressed voice command; and performing audio enhancement on the suppressed voice command to obtain a voice command.

[0010] Optionally, inputting the voice command into the speech recognition model and processing it to obtain the speech recognition text includes: the speech recognition model cutting the voice command according to the configured sliding window parameters to obtain N voice segments, where N is a positive integer; for a voice segment, the speech recognition model identifying whether there are keywords in the keyword list in the voice segment; if there are keywords in the keyword list in the voice segment, the speech recognition model determining the frequency of occurrence of the keywords in the voice segment; obtaining frequency rules, retaining the keywords in the voice segment if the frequency matches the frequency rules, and deleting the keywords in the voice segment if the frequency does not match the frequency rules; obtaining the retained keywords from the N voice segments to obtain Y keywords, and obtaining the N voice segment-related interjections, combining the Y keywords and interjections to obtain the speech recognition text, where Y is less than or equal to N and Y is a positive integer.

[0011] Optionally, the speech recognition model is obtained as follows: M historical speech commands within a historical time period are obtained, and the historical speech recognition text corresponding to each historical speech command is obtained, where M is a positive integer; the preset speech recognition model is trained using the M historical speech commands and the M historical speech recognition text to obtain the trained speech recognition model; sliding window parameters and a keyword list are obtained, and the parameters of the trained speech recognition model are configured using the sliding window parameters and the keyword list to obtain the speech recognition model.

[0012] Optionally, determining the financial business type based on the speech recognition text includes: obtaining Y keywords associated with the speech recognition text, and obtaining the keyword type of the keyword in a preset position among the Y keywords to obtain an initial keyword type, where Y is a positive integer; if the initial keyword type is a wake-up type, determining the financial business type as a voice acquisition type; if the initial keyword type is a business execution type, obtaining the keyword types of keywords other than the keyword in the preset position to obtain K target keyword types, where K is a positive integer; and determining the financial business type based on the K target keyword types.

[0013] Optionally, determining the financial business type based on the K target keyword types includes: obtaining a list of corresponding keywords, wherein the list of corresponding keywords includes multiple keywords and the business type corresponding to each keyword, and the business type includes at least one of the following: process business type, query business type, detection business type, and add / delete business type; extracting K business types from the list of corresponding keywords based on the K target keyword types, and constructing the financial business type from the K business types.

[0014] To achieve the above objectives, according to another aspect of this application, a financial transaction processing apparatus is provided. The apparatus includes: a receiving unit for receiving a voice command sent by a target user through a client, wherein the voice command instructs the target user to apply for a financial transaction; an input unit for inputting the voice command into a speech recognition model and processing it to obtain speech recognition text, wherein the speech recognition model converts the voice command to obtain speech recognition text; a determining unit for determining the type of financial transaction based on the speech recognition text; and a calling unit for calling a target business model according to the type of financial transaction and controlling the target business model to execute the financial transaction associated with the voice command.

[0015] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform any of the above-mentioned financial business processing methods.

[0016] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory storing an executable program, and the processor for running the program, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-described financial business processing methods.

[0017] According to another aspect of the present invention, a computer program product is also provided, the computer program product including a computer program, wherein when the computer program is executed by a processor, it implements the processing method of any of the above-mentioned financial transactions.

[0018] In this embodiment, a financial transaction processing method is adopted. This involves receiving voice commands sent by a target user through a client, where the voice commands instruct the target user on the financial transaction they wish to apply for. The voice commands are then input into a speech recognition model to obtain speech-recognized text. The speech recognition model converts the voice commands into speech-recognized text. The type of financial transaction is determined based on the speech-recognized text. A target business model is then invoked based on the financial transaction type, and the target business model is controlled to execute the financial transaction associated with the voice command. This method solves the technical problems of low speech recognition accuracy and low business processing efficiency in related technologies when processing financial transactions via voice commands. By inputting voice commands into a speech recognition model, obtaining speech-recognized text, invoking the target business model based on the financial transaction type, and controlling the target business model to execute the financial transaction associated with the voice command, the technical effect of improving the speech recognition accuracy and business processing efficiency of processing financial transactions via voice commands is achieved. Attached Figure Description

[0019] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 It is a hardware structure block diagram of a computer terminal (or mobile device) used to implement a processing method for financial business.

[0021] Figure 2 This is a flowchart of a financial transaction processing method provided according to an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of a financial transaction processing system provided according to an embodiment of this application;

[0023] Figure 4 This is a flowchart of a speech processing method provided according to an embodiment of this application;

[0024] Figure 5This is a schematic diagram of a business process reminder system provided according to an embodiment of this application;

[0025] Figure 6 This is a schematic diagram of a financial transaction processing apparatus provided according to an embodiment of this application;

[0026] Figure 7 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties. For example, this system has interfaces with relevant users or organizations to provide users with corresponding operation data for them to choose to agree to or refuse automated decision-making results. Before obtaining relevant information, a request for obtaining the information needs to be sent to the aforementioned user or organization through the interface, and the relevant information is obtained after receiving consent from the aforementioned user or organization; if the user chooses to refuse, the expert decision-making process is initiated. Users can view the purpose of data use in real time through authorization decoding and have the right to withdraw authorization or delete data at any time. After the authorization is withdrawn, the system will terminate the relevant data processing within 24 hours.

[0030] It should be noted that the information collected in this application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize use or refuse use.

[0031] Example 1

[0032] According to an embodiment of this application, a method embodiment for processing financial transactions is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) used to implement financial business processing methods, such as... Figure 1 As shown, computer terminal 10 (or mobile device) may include one or more ( Figure 1 The processor 102 (which may include, but is not limited to, a microprocessor MCU (Microcontroller Unit) or a programmable gate array (FPGA)) is shown as 102a, 102b, ..., 102n. It also includes a memory 104 for storing data and a transmission device 106 for communication functions. In addition, it may include: a display, an input / output interface, a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a keyboard, a cursor control device, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0034] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the financial business processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned financial business processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) and a network interface, which can be connected to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a radio frequency (RF) module, used for wireless communication with the Internet.

[0037] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0038] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for processing financial transactions is shown. Figure 2 This is a flowchart of a financial transaction processing method provided according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0039] Step S201: Receive a voice command sent by the target user through the client, wherein the voice command instructs the target user to apply for a financial service.

[0040] It should be noted that the target users can refer to staff such as counter workers at financial institutions and project commissioning coordinators, while the client refers to the interactive client deployed on the counter terminals and project command screens of financial institutions. When the target user sends a dictated voice message through the client, the client can receive voice commands sent through the client, which can instruct the target user on the financial transaction they are applying for.

[0041] Step S202: Input the voice command into the voice recognition model and process it to obtain the voice recognition text. The voice recognition model is used to convert the voice command to obtain the voice recognition text.

[0042] It should be noted that the speech recognition model can be a dual-mode combination model consisting of a keyword detection tool and a speech recognition tool. It can process the voice commands sent by the target user and output speech-recognized text. The keyword detection tool can quickly scan the audio to detect whether there are keywords. If no keywords are found, they are discarded directly. If they are found, the speech recognition tool is used to convert the text.

[0043] Step S203: Determine the type of financial business based on the speech-recognized text.

[0044] Specifically, after obtaining the speech recognition text output by the speech recognition model, the content of the text can be subjected to fuzzy semantic understanding and intent recognition to determine the financial business type (i.e. intent type). The financial business type refers to the business process category label set by the financial institution. Each label corresponds to a manageable and structured business process template, such as business query, step addition, step completion mark, etc.

[0045] Step S204: Invoke the target business model according to the type of financial business, and control the target business model to execute the financial business associated with the voice command.

[0046] Specifically, after obtaining the type of financial business, the corresponding business model can be determined based on that type, i.e., the target business model. The target business model refers to the logical processing module corresponding to different business intentions, which can connect to business retrieval interfaces, business editing interfaces, process status update interfaces, etc. After determining the target business model, it can be invoked, and the model can be controlled to execute the business process operations corresponding to voice commands, and the results can be fed back to the target user.

[0047] The financial transaction processing method provided in this application embodiment receives a voice command sent by a target user through a client, wherein the voice command instructs the target user to apply for a financial transaction; inputs the voice command into a voice recognition model to obtain voice-recognized text, wherein the voice recognition model is used to convert the voice command into voice-recognized text; determines the type of financial transaction based on the voice-recognized text; and calls a target business model based on the type of financial transaction and controls the target business model to execute the financial transaction associated with the voice command. This method solves the technical problems of low voice recognition accuracy and low business processing efficiency in related technologies when processing financial transactions via voice commands. By inputting the voice command into a voice recognition model, processing it to obtain voice-recognized text, calling the target business model based on the type of financial transaction, and controlling the target business model to execute the financial transaction associated with the voice command, the method achieves the technical effect of improving the voice recognition accuracy and business processing efficiency of processing financial transactions via voice commands.

[0048] Optionally, in the financial business processing method provided in the embodiments of this application, before receiving the voice command sent by the target user through the client, the method further includes: detecting whether the target user generates a trigger command through the client, wherein the trigger command is generated when the target user triggers a preset wake word through the client, or when the target user is in a preset scenario; if the trigger command is detected to be generated by the target user through the client, the step of receiving the voice command sent by the target user through the client is executed.

[0049] Specifically, before the target user sends a voice command through the client, the system first needs to detect the target user's operation or environmental status to determine whether the conditions for starting voice interaction are met, that is, to detect whether the target user generates a trigger command through the client. The trigger command, as a start signal, can be generated by the target user through the client by triggering a preset wake word. For example, the preset wake word can be "process assistant". It can also be generated by the client when the target user is in a preset scenario. For example, the target user can directly generate a trigger command by clicking the "voice input" button on the physical equipment of a financial institution branch.

[0050] Furthermore, when a trigger command is detected generated by the target user through the client, that is, after the trigger command is detected, the subsequent steps of receiving voice commands sent by the target user through the client can be executed, avoiding the burden, misidentification and privacy risks caused by continuous monitoring.

[0051] This embodiment detects trigger commands and maintains voice silence when there is no user-initiated operation or when the relevant scenario is not entered. This reduces resource consumption, minimizes the risk of misidentification, protects data privacy, and achieves precise synchronization between voice interaction and business processes.

[0052] Optionally, in the financial business processing method provided in this application embodiment, receiving a voice command sent by a target user through a client includes: receiving an initial voice command sent by the target user through a client with the authorization of the target user; performing silence detection on the initial voice command; and if there is noise in the initial voice command, performing noise suppression on the initial voice command to obtain a suppressed voice command; and performing audio enhancement on the suppressed voice command to obtain a voice command.

[0053] Specifically, to acquire voice commands and avoid continuous monitoring without the user's knowledge or consent, thus preventing the accidental recording or transmission of sensitive business voice messages, the initial voice command sent by the target user can be acquired first, with the target user's authorization. Then, silence detection is performed on this command, that is, determining whether the audio of the acquired initial voice command contains invalid speech, such as air noise or background noise. This can be identified using algorithms such as energy thresholding or zero-crossing rate.

[0054] If noise is present in the initial voice command, it is necessary to eliminate the interference of noise on speech recognition, improve the subsequent recognition accuracy, and avoid misjudgment of keywords or text transcription errors due to background noise. This means that noise suppression is performed on the initial voice command to obtain the suppressed voice command. Then, audio enhancement is performed on the suppressed voice command, which means further optimization of the noise-suppressed voice signal. This can include processing such as volume normalization, spectrum compensation, and speech rate smoothing to make the voice features more consistent with the input requirements of the speech recognition model and improve the clarity of low-volume commands, thereby obtaining the final voice command.

[0055] This embodiment effectively improves the signal-to-noise ratio and recognition reliability of voice input by performing silence detection, noise suppression, and audio enhancement processing on voice commands, ensuring accurate transcription of voice commands in complex environments and providing a stable and high-quality voice data foundation for subsequent intent parsing and process execution.

[0056] Optionally, in the financial business processing method provided in this application embodiment, inputting a voice command into a voice recognition model and processing it to obtain voice recognition text includes: the voice recognition model cutting the voice command according to configured sliding window parameters to obtain N voice segments, where N is a positive integer; for a voice segment, the voice recognition model identifying whether there are keywords in the keyword list in the voice segment; if there are keywords in the keyword list in the voice segment, the voice recognition model determining the frequency of occurrence of the keywords in the voice segment; obtaining frequency rules, retaining the keywords in the voice segment if the frequency matches the frequency rules, and deleting the keywords in the voice segment if the frequency does not match the frequency rules; obtaining the retained keywords from the N voice segments to obtain Y keywords, and obtaining the N voice segments associated with interjections, combining the Y keywords and interjections to obtain voice recognition text, where Y is less than or equal to N and Y is a positive integer.

[0057] It should be noted that the sliding window parameters refer to the preset time length and step size, such as extracting a segment of speech every 200 milliseconds and moving it forward by 100 milliseconds each time to form overlapping segments. In order to facilitate segment-by-segment processing by the model, improve the local accuracy of keyword detection, avoid misjudgment due to excessively long speech, and adapt to the characteristics of uneven speech rate and frequent pauses in the speech of financial institutions, after the speech recognition model obtains the speech command, it can use the above sliding window parameters to cut the speech command to obtain multiple short-time speech units, that is, multiple speech segments.

[0058] Furthermore, after obtaining the speech segments, the speech recognition model identifies whether each speech segment contains keywords from the keyword list, quickly filtering out speech segments that may contain business instructions, skipping meaningless segments, reducing the burden of subsequent processing, and improving overall response efficiency. The keyword list contains a financial institution's exclusive vocabulary, which includes keywords for high-frequency operations such as query, add, and start check.

[0059] For a given speech segment, if a keyword from the keyword list appears within that segment, the speech recognition model can determine the frequency of that keyword's occurrence within the current window, and use this frequency for verification. If the frequency matches a frequency rule—for example, if the rule is that a keyword's frequency in three consecutive speech segments is greater than or equal to 1 and less than or equal to 2, and there are no other keywords competing within the preceding or following second—then that keyword can be retained. Conversely, if the frequency is too low, it may be due to noise and false triggering, and the keyword should be discarded. It should be noted that if multiple keywords appear with high probability simultaneously (e.g., "query" and "add" appear simultaneously), the longest match priority can be used, prioritizing longer keywords (e.g., "add core steps" is longer than "add"). If the same keyword is repeatedly triggered within 10 seconds, this operation is prone to error; therefore, repeated wake-ups can be suppressed.

[0060] Furthermore, the keywords retained from all the above speech segments are obtained, as well as the interjections associated with all speech segments. Then, the keywords and interjections are combined according to the time sequence of the speech segments, and the retained keywords are concatenated with the context interjections to form a complete and natural recognition text, thus obtaining the speech recognition text.

[0061] This embodiment achieves accurate screening and semantic reconstruction of voice commands by using sliding window segmentation, keyword existence judgment, frequency filtering, and interjection fusion, effectively improving the accuracy and semantic integrity of speech recognition text and reducing the risk of misoperation.

[0062] Optionally, in the financial business processing method provided in this application embodiment, the speech recognition model is obtained in the following way: obtaining M historical voice commands within a historical time period, and obtaining the historical speech recognition text corresponding to each historical voice command, where M is a positive integer; training a preset speech recognition model using the M historical voice commands and M historical speech recognition text to obtain a trained speech recognition model; obtaining sliding window parameters and a keyword list, and configuring the parameters of the trained speech recognition model using the sliding window parameters and the keyword list to obtain a speech recognition model.

[0063] Specifically, to accurately recognize voice commands, the speech recognition model first needs to be trained. This begins by acquiring historical voice commands from multiple time periods, followed by their corresponding historical speech-recognition texts. This provides the speech recognition model with authentic and accurate voice-text pairing data, enabling it to learn pronunciation patterns, terminology, and pragmatic habits specific to financial institutions, thus improving its understanding of financial business speech. The pre-set speech recognition model is then trained using the historical voice commands and their corresponding historical speech-recognition texts. This pre-set model refers to an initial, lightweight, open-source speech recognition model with basic speech-to-text capabilities, but not optimized for financial institution business.

[0064] Furthermore, the sliding window parameters and keyword list are obtained, and then the parameters of the trained speech recognition model are configured using the sliding window parameters and keyword list. That is, the above two external configuration items are bound to the trained model, so that it pays priority to these keywords during runtime, uses a specified window for segmentation processing, does not change the model structure, only adjusts the behavior logic, and thus obtains the speech recognition model.

[0065] This embodiment trains the basic model using historical speech and text data from real financial institution scenarios, and combines it with business-specific sliding windows and keyword lists for parameter binding, forming a lightweight, offline, and accurate dedicated speech recognition model. This model does not require an internet connection or rely on the cloud, and can stably recognize business instructions in the closed environment of financial institutions. It effectively solves the problems of poor recognition of financial terms and high misrecognition rate in noisy environments of general speech models, realizing the practical deployment of voice input.

[0066] Optionally, in the financial business processing method provided in this application embodiment, determining the financial business type based on the speech recognition text includes: obtaining Y keywords associated with the speech recognition text, and obtaining the keyword type of the keyword in a preset position among the Y keywords to obtain an initial keyword type, where Y is a positive integer; if the initial keyword type is a wake-up type, determining the financial business type as a voice acquisition type; if the initial keyword type is a business execution type, obtaining the keyword types of keywords other than the keyword in the preset position to obtain K target keyword types, where K is a positive integer; and determining the financial business type based on the K target keyword types.

[0067] Specifically, after obtaining the speech recognition text output by the model, the system first extracts multiple keywords from the speech recognition text that belong to the preset keyword list, and then extracts the keyword type of the keyword in a fixed order position in the text, such as the keyword type of the first word, to obtain the initial keyword type. This allows the system to determine whether the user's current intent is to activate the voice function or to initiate a specific business operation, thus achieving a preliminary classification of intent and reducing the burden of subsequent processing.

[0068] If the initial keyword type is a wake-up type, which refers to keywords used to activate voice interaction functions, such as a workflow assistant, then the financial business type can be determined as the voice acquisition type. If the initial keyword type is a business execution type, that is, explicitly pointing to a specific operation, then it is necessary to obtain the keyword type of keywords other than those in the preset positions. For example, if the query is a business query type, by extracting the business attributes from subsequent keywords, the specific business content can be accurately identified. In other words, the financial business type (i.e., the intent type) can be determined based on these target keyword types.

[0069] This embodiment avoids misjudgments caused by relying solely on a single word or location by judging the keyword type, ensuring that voice commands are correctly mapped to the corresponding business processes, thereby improving the accuracy of intent recognition and operational reliability for financial institutions in multiple scenarios.

[0070] Optionally, in the financial business processing method provided in this application embodiment, determining the financial business type based on K target keyword types includes: obtaining a keyword correspondence list, wherein the keyword correspondence list includes multiple keywords and the business type corresponding to each keyword, and the business type includes at least one of the following: process business type, query business type, detection business type, and add / delete business type; extracting K business types from the keyword correspondence list based on the K target keyword types, and constructing the K business types into a financial business type.

[0071] Specifically, after obtaining multiple target keyword types, the first step is to acquire a list of corresponding keywords. This list records the association between each keyword and its associated business type. For example, "query" corresponds to the "query" business type, "add" to the "add / delete" business type, "process" business type refers to a complete operation process that requires multiple steps to be executed sequentially, "query" business type refers to operations used only to obtain information without making changes, "detection" business type refers to operations used to trigger checks and verifications, and "add / delete" business type refers to operations involving adding, deleting, or modifying configurations. Then, using the target keyword types, the corresponding business types are extracted from the keyword list, and these business types are combined to form financial business types. For example, financial business types include "add / delete" business types and "process" business types. At this point, entries containing both "add" actions and "process" business types can be searched in the process library, and the corresponding steps can be matched.

[0072] This embodiment establishes a static mapping list between keywords and business types, and uses multiple target keyword types for combination matching. This can transform scattered keyword intents into complete financial business types, improving the parsing accuracy and execution accuracy of voice commands in complex business scenarios.

[0073] This application also provides a financial transaction processing system. Figure 3 This is a schematic diagram of a financial transaction processing system provided according to an embodiment of this application, such as... Figure 3 As shown, the system includes: a core functional module, a voice interaction module, a data storage module, and an adaptation layer module. The core functional module includes a process management unit, a process reminder unit, and a data management unit. The voice interaction module includes a voice input unit, a voice recognition and processing unit, an intent parsing unit, and a voice synthesis unit, enabling data interaction with the user. The data storage module stores business process data, record data, and voice interaction configuration data. The adaptation layer module enables interface adaptation between the voice interaction module and the core functional module.

[0074] Figure 4 This is a flowchart of a speech processing method provided according to an embodiment of this application, such as... Figure 4 As shown, at the start of the voice interaction process, it can first determine whether the target user has activated the voice interaction module. In this case, the target user can activate the voice interaction function using a preset wake-up word or by directly triggering a specific voice operation scenario. If not activated, it remains in standby mode. When the voice interaction module is activated, the voice input unit first performs voice acquisition and noise filtering, i.e., acquiring the target user's initial voice command, processing it through a noise filtering algorithm, and then transmitting the voice command to the voice recognition processing unit. The voice recognition processing unit locates keywords in the speech associated with the initial voice command, converts the speech associated with the initial voice command into text, and then uses Chinese word segmentation and part-of-speech tagging methods to perform text error correction, generating standardized speech recognition text.

[0075] Furthermore, the intent parsing unit performs intent recognition on the speech-recognized text, which can be achieved by matching intent types in the intent rule base. When determining the intent type, if it is a business query intent, the core functional module's business retrieval interface is called, returning the matching business process. The query result is then read aloud by the speech synthesis unit, and the process details are displayed on the interface. If it is a step entry intent, the business editing interface is called. If it is a process control intent, the process execution interface is called to execute the operation and provide feedback. If it is a reminder configuration intent, the system reminder parameters are updated. After executing the corresponding operation, the speech synthesis unit performs a speech synthesis and playback operation, and relevant information is simultaneously displayed on the business process reminder system interface, ending the process. It should be noted that... Figure 5 This is a schematic diagram of a business process reminder system provided according to an embodiment of this application, such as... Figure 5 As shown, the system interface of the business process reminder system can include a voice interaction area, a voice broadcast area, and a process display area. The voice interaction area can include multiple functional components such as voice interaction entry, completion marker, initiation of compliance check, and process settings. The process display area can display the steps to be completed and the steps that have been completed. After the operation is completed, a log is recorded and stored locally, the interaction ends, and a closed-loop management is formed.

[0076] This embodiment improves the accuracy of speech recognition and the efficiency of processing financial transactions by inputting voice commands into a speech recognition model, processing the voice commands into speech recognition text, calling the target business model according to the type of financial transaction, and controlling the target business model to execute the financial transaction associated with the voice command.

[0077] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0078] Example 2

[0079] This application also provides a financial transaction processing apparatus. It should be noted that the financial transaction processing apparatus of this application can be used to execute the financial transaction processing method provided in this application. The following describes the financial transaction processing apparatus provided in this application.

[0080] According to an embodiment of this application, an apparatus for implementing the above-described financial business processing method is also provided. Figure 6 This is a schematic diagram of a financial transaction processing apparatus provided according to an embodiment of this application, such as... Figure 6As shown, the device includes: a receiving unit 60, an input unit 61, a determining unit 62, and a calling unit 63.

[0081] The receiving unit 60 is used to receive voice commands sent by the target user through the client, wherein the voice commands indicate the financial business applied for by the target user;

[0082] The input unit 61 is used to input voice commands into the speech recognition model and process them to obtain speech recognition text. The speech recognition model is used to convert the voice commands to obtain speech recognition text.

[0083] Determining unit 62 is used to determine the type of financial business based on the speech recognition text;

[0084] Calling unit 63 is used to call the target business model according to the financial business type and control the target business model to execute the financial business associated with the voice command.

[0085] The financial transaction processing apparatus provided in this application embodiment receives voice commands sent by a target user through a client via a receiving unit 60. The voice commands indicate the financial transaction the target user is applying for. An input unit 61 inputs the voice commands into a voice recognition model, which processes them to obtain voice-recognized text. The voice recognition model converts the voice commands into voice-recognized text. A determining unit 62 determines the type of financial transaction based on the voice-recognized text. A calling unit 63 calls a target business model based on the type of financial transaction and controls the target business model to execute the financial transaction associated with the voice command. This solves the technical problems of low voice recognition accuracy and low business processing efficiency in related technologies when processing financial transactions via voice commands. By inputting voice commands into a voice recognition model, processing them to obtain voice-recognized text, calling the target business model based on the type of financial transaction, and controlling the target business model to execute the financial transaction associated with the voice command, the apparatus achieves the technical effect of improving the voice recognition accuracy and business processing efficiency of processing financial transactions via voice commands.

[0086] Optionally, in the financial business processing apparatus provided in this application embodiment, the apparatus further includes: a detection unit, configured to detect whether the target user generates a trigger command through the client before receiving a voice command sent by the target user through the client, wherein the trigger command is generated when the target user triggers a preset wake word through the client, or when the target user is in a preset scenario; and an execution unit, configured to execute the step of receiving the voice command sent by the target user through the client when the trigger command is detected.

[0087] Optionally, in the financial business processing apparatus provided in this application embodiment, the receiving unit 60 includes: a receiving module, configured to receive an initial voice command sent by the target user through a client when authorized by the target user; a detection module, configured to perform silence detection on the initial voice command, and suppress noise in the initial voice command if noise exists, to obtain a suppressed voice command; and an enhancement module, configured to enhance the suppressed voice command to obtain a voice command.

[0088] Optionally, in the financial business processing apparatus provided in this application embodiment, the input unit 61 includes: a cutting module, used to cut the voice command by the speech recognition model according to the configured sliding window parameters to obtain N voice segments, where N is a positive integer; a recognition module, used to identify whether a keyword in the keyword list exists in a voice segment for a given voice segment; a first determination module, used to determine the frequency of occurrence of the keyword in the voice segment by the speech recognition model when the keyword in the keyword list exists in the voice segment; a first acquisition module, used to acquire frequency rules, retain the keyword in the voice segment when the frequency matches the frequency rules, and delete the keyword in the voice segment when the frequency does not match the frequency rules; and a second acquisition module, used to acquire the retained keywords in the N voice segments to obtain Y keywords, and acquire the N voice segments associated with interjections, and combine the Y keywords and interjections to obtain the speech recognition text, where Y is less than or equal to N and Y is a positive integer.

[0089] Optionally, in the financial business processing apparatus provided in this application embodiment, the input unit 61 includes: a third acquisition module, used to acquire M historical voice commands within a historical time period, and acquire the historical speech recognition text corresponding to each historical voice command, where M is a positive integer; a training module, used to train a preset speech recognition model using the M historical voice commands and the M historical speech recognition text to obtain a trained speech recognition model; and a fourth acquisition module, used to acquire sliding window parameters and a keyword list, and to configure the parameters of the trained speech recognition model using the sliding window parameters and the keyword list to obtain a speech recognition model.

[0090] Optionally, in the financial business processing apparatus provided in this application embodiment, the determining unit 62 includes: a fifth acquisition module, used to acquire Y keywords associated with the speech recognition text, and acquire the keyword type of the keyword in the preset position among the Y keywords to obtain an initial keyword type, where Y is a positive integer; a second determining module, used to determine the financial business type as a voice acquisition type when the initial keyword type is a wake-up type; a sixth acquisition module, used to acquire the keyword type of keywords other than the keyword in the preset position when the initial keyword type is a business execution type, to obtain K target keyword types, where K is a positive integer; and a third determining module, used to determine the financial business type based on the K target keyword types.

[0091] Optionally, in the financial business processing apparatus provided in this application embodiment, the determining unit 62 includes: a seventh acquisition module, used to acquire a keyword correspondence list, wherein the keyword correspondence list includes multiple keywords and a business type corresponding to each keyword, and the business type includes at least one of the following: process business type, query business type, detection business type, and add / delete business type; and an extraction module, used to extract K business types from the keyword correspondence list according to K target keyword types, and to form a financial business type from the K business types.

[0092] It should be noted that the receiving unit 60, input unit 61, determining unit 62, and calling unit 63 mentioned above correspond to steps S201 to S204 in Embodiment 1. The instances and application scenarios implemented by the above units and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.

[0093] Example 3

[0094] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced with a mobile terminal or an electronic device, etc.

[0095] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0096] In this embodiment, the computer terminal described above can execute the following steps in the financial business processing method: receiving a voice command sent by a target user through a client, wherein the voice command indicates the financial business applied for by the target user; inputting the voice command into a voice recognition model and processing it to obtain voice recognition text, wherein the voice recognition model is used to convert the voice command to obtain voice recognition text; determining the financial business type based on the voice recognition text; calling the target business model based on the financial business type and controlling the target business model to execute the financial business associated with the voice command.

[0097] Optionally, the computer terminal described above can execute the program code for the following steps in the financial business processing method: detecting whether the target user generates a trigger command through the client, wherein the trigger command is generated when the target user triggers a preset wake word through the client, or when the target user is in a preset scenario; and if the trigger command is detected to be generated by the target user through the client, executing the step of receiving the voice command sent by the target user through the client.

[0098] Optionally, the computer terminal described above can execute the program code for the following steps in the financial business processing method: receiving an initial voice command sent by the target user through a client when authorized by the target user; performing silence detection on the initial voice command; and, if noise exists in the initial voice command, suppressing the noise to obtain a suppressed voice command; and enhancing the audio of the suppressed voice command to obtain a voice command.

[0099] Optionally, the computer terminal described above can execute the following steps in the financial business processing method: The speech recognition model cuts the speech command according to the configured sliding window parameters to obtain N speech segments, where N is a positive integer; for each speech segment, the speech recognition model identifies whether a keyword from the keyword list exists in the speech segment; if a keyword from the keyword list exists in the speech segment, the speech recognition model determines the frequency of occurrence of the keyword in the speech segment; frequency rules are obtained, and if the frequency matches the frequency rule, the keyword in the speech segment is retained; if the frequency does not match the frequency rule, the keyword in the speech segment is deleted; the retained keywords from the N speech segments are obtained to obtain Y keywords, and N voice words associated with the speech segments are obtained; the Y keywords and voice words are combined to obtain the speech recognition text, where Y is less than or equal to N and Y is a positive integer.

[0100] Optionally, the computer terminal described above can execute the following steps in the financial business processing method: obtaining M historical voice commands within a historical time period, and obtaining the historical speech recognition text corresponding to each historical voice command, where M is a positive integer; training a preset speech recognition model using the M historical voice commands and M historical speech recognition texts to obtain the trained speech recognition model; obtaining sliding window parameters and a keyword list, and configuring the parameters of the trained speech recognition model using the sliding window parameters and the keyword list to obtain the speech recognition model.

[0101] Optionally, the computer terminal described above can execute the following steps in the financial business processing method: obtaining Y keywords associated with the speech recognition text, and obtaining the keyword type of the keyword in the preset position among the Y keywords to obtain the initial keyword type, where Y is a positive integer; if the initial keyword type is a wake-up type, determining the financial business type as a voice acquisition type; if the initial keyword type is a business execution type, obtaining the keyword types of keywords other than the keyword in the preset position to obtain K target keyword types, where K is a positive integer; determining the financial business type based on the K target keyword types.

[0102] Optionally, the computer terminal described above can execute the program code for the following steps in the financial business processing method: obtaining a keyword correspondence list, wherein the keyword correspondence list includes multiple keywords and the business type corresponding to each keyword, and the business type includes at least one of the following: process business type, query business type, detection business type, and add / delete business type; extracting K business types from the keyword correspondence list based on K target keyword types, and constructing the K business types into a financial business type.

[0103] Optionally, Figure 7 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 7 As shown, the electronic device may include: one or more ( Figure 7 (Only one is shown) processor 702, memory 704, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0104] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the financial business processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned financial business processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0105] The processor can access the information and application programs stored in the memory via the transmission device to execute the steps described above in the financial business processing method.

[0106] Those skilled in the art will understand that Figure 7 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 7 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 7 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 7 The different configurations shown.

[0107] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0108] Example 4

[0109] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the financial transaction processing method provided in Embodiment 1.

[0110] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0111] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a voice command sent by a target user through a client, wherein the voice command indicates the financial service applied for by the target user; inputting the voice command into a voice recognition model and processing it to obtain voice recognition text, wherein the voice recognition model is used to convert the voice command to obtain voice recognition text; determining the type of financial service based on the voice recognition text; invoking a target business model based on the type of financial service and controlling the target business model to execute the financial service associated with the voice command.

[0112] This application also provides a computer program product, which, when executed on a data processing device, is suitable for performing the processing steps of a financial transaction.

[0113] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0114] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0115] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0116] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0117] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0118] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0119] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for processing financial transactions, characterized in that, include: Receive voice commands sent by a target user through a client, wherein the voice commands indicate the financial services applied for by the target user; The voice command is input into a speech recognition model and processed to obtain speech recognition text, wherein the speech recognition model is used to convert the voice command to obtain the speech recognition text; The type of financial transaction is determined based on the speech-recognized text. The target business model is invoked according to the financial business type, and the target business model is controlled to execute the financial business associated with the voice command.

2. The method according to claim 1, characterized in that, Before receiving voice commands sent by the target user through the client, the method further includes: Detect whether the target user generates a trigger command through the client, wherein the trigger command is generated when the target user triggers a preset wake word through the client, or when the target user is in a preset scenario; If it is detected that the target user generates the trigger command through the client, the step of receiving the voice command sent by the target user through the client is executed.

3. The method according to claim 1, characterized in that, Receiving voice commands sent by the target user through the client includes: With the authorization of the target user, receive the initial voice command sent by the target user through the client; The initial voice command is subjected to silence detection. If noise is present in the initial voice command, the noise is suppressed to obtain the suppressed voice command. The suppressed speech command is then enhanced to obtain the speech command.

4. The method according to claim 1, characterized in that, The voice command is input into the speech recognition model, and the resulting speech recognition text includes: The speech recognition model cuts the speech command according to the configured sliding window parameters to obtain N speech segments, where N is a positive integer; For a given speech segment, the speech recognition model identifies whether a keyword from a keyword list exists in the speech segment. If a keyword from the keyword list exists in the speech segment, the speech recognition model determines the frequency of occurrence of the keyword in the speech segment. Obtain frequency rules; if the frequency of occurrence matches the frequency rule, retain the keywords in the speech segment; if the frequency of occurrence does not match the frequency rule, delete the keywords in the speech segment. The keywords retained in the N speech segments are obtained to obtain Y keywords, and the tone words associated with the N speech segments are obtained. The Y keywords and the tone words are combined to obtain the speech recognition text, where Y is less than or equal to N and Y is a positive integer.

5. The method according to claim 4, characterized in that, The speech recognition model is obtained in the following way: Get M historical voice commands within a historical time period, and get the historical speech recognition text corresponding to each historical voice command, where M is a positive integer; The preset speech recognition model is trained using the M historical speech commands and M historical speech recognition texts to obtain the trained speech recognition model. Obtain the sliding window parameters and the keyword list, and configure the parameters of the trained speech recognition model using the sliding window parameters and the keyword list to obtain the speech recognition model.

6. The method according to claim 1, characterized in that, The financial transaction type determined based on the speech-recognized text includes: Obtain Y keywords associated with the speech-recognized text, and obtain the keyword type of the keyword in a preset position among the Y keywords to obtain the initial keyword type, where Y is a positive integer; If the initial keyword type is a wake-up type, the financial business type is determined to be a voice acquisition type; If the initial keyword type is a business execution type, obtain the keyword types of keywords other than the keywords in the preset position to get K target keyword types, where K is a positive integer; The financial business type is determined based on the K target keyword types.

7. The method according to claim 6, characterized in that, The financial business type is determined based on the K target keyword types, including: Obtain a list of corresponding keywords, wherein the list includes multiple keywords and the business type corresponding to each keyword, and the business type includes at least one of the following: process business type, query business type, detection business type, and add / delete business type; Based on the K target keyword types, K business types are extracted from the keyword corresponding list, and the K business types constitute the financial business type.

8. A financial transaction processing device, characterized in that, include: A receiving unit is configured to receive voice commands sent by a target user through a client, wherein the voice commands indicate the financial services applied for by the target user; An input unit is used to input the voice command into a speech recognition model and process it to obtain speech recognition text, wherein the speech recognition model is used to convert the voice command to obtain the speech recognition text; The determining unit is used to determine the type of financial transaction based on the speech-recognized text; The invocation unit is used to invoke the target business model according to the financial business type, and control the target business model to execute the financial business associated with the voice command.

9. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the financial transaction processing method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the financial business processing method according to any one of claims 1 to 7.