Cross-border financial transaction control method and device, computer equipment and storage medium
By building a multilingual voiceprint feature model and a dynamic threshold mechanism, the deficiencies in multilingual processing and security strategies in cross-border financial transactions are addressed, efficient and accurate voiceprint authentication is achieved, and the security and user experience of cross-border transactions are improved.
Patent Information
- Application Number
- CN202510671638.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-09
AI Technical Summary
The existing voiceprint authentication system cannot effectively handle multilingual scenarios in cross-border financial transactions, resulting in low recognition accuracy, high verification failure rate, inability to dynamically adjust security levels and low resource utilization efficiency, and cannot meet the real-time and security requirements of cross-border financial transactions.
Build a multilingual voiceprint feature model, extract cross-language common features and language-specific features through shared encoders and adapters, dynamically adjust the voiceprint matching threshold based on scene perception information, achieve accurate segmentation and weighted fusion of multiple languages, support mixed input of multiple languages such as Chinese, English, Cantonese, and improve verification strictness in high-risk transactions.
It improves the accuracy and success rate of voiceprint verification, reduces the false acceptance rate, reduces resource waste, improves system response speed and user experience, and enhances the security and efficiency of cross-border financial transactions.
Smart Images

Figure CN120612945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial transaction technology, and more specifically to a cross-border financial transaction control method, apparatus, computer equipment, and storage medium. Background Art
[0002] Against the backdrop of accelerating global economic integration, cross-border financial transactions are becoming increasingly frequent and their scale continues to expand. Users are placing higher demands on the security, convenience, and efficiency of these transactions. Voiceprint authentication, as an emerging biometric technology, offers broad application prospects in cross-border financial transactions thanks to its unique advantages in authentication, such as difficulty in forgery, convenience, and efficiency. However, the voiceprint authentication systems currently in widespread use in cross-border financial transactions face numerous technical challenges that need to be addressed.
[0003] Currently, most voiceprint authentication systems are primarily built for a single language (such as Chinese). While this has optimized the voiceprint registration process to a certain extent and supports Mandarin verification, it is insufficient when faced with the multilingual mixed-language scenarios of users in cross-border transactions. Cross-border transactions involve users from different countries and regions, with diverse language habits and accents. For example, mixed Chinese and English communication and dialect accents are common. Existing single-language models are unable to effectively cover these multilingual requirements. As a result, during the actual verification process, the system's recognition accuracy for such complex speech is reduced, and the verification failure rate is significantly increased, seriously affecting user experience and transaction efficiency.
[0004] Furthermore, while some voiceprint authentication systems claim to support multiple languages, they exhibit significant limitations in practice. These systems require users to preselect a language, and model switching relies on manual configuration. In scenarios requiring extremely high real-time performance, such as cross-border financial transactions, manual model configuration is not only cumbersome and time-consuming, but also fails to meet real-time requirements. When users temporarily switch languages during a transaction, the system is unable to respond promptly, disrupting the transaction process and potentially incurring transaction risks.
[0005] Furthermore, traditional voiceprint authentication systems employ a static threshold strategy, using a fixed voiceprint matching threshold for identity verification. However, cross-border financial transactions are complex and diverse, and the security risks faced by different scenarios vary. For example, high-risk cross-border transfers require a higher security level, while standard query transactions have relatively lower security requirements. Fixed threshold strategies cannot dynamically adjust security levels based on actual transaction scenarios. In high-risk transactions, insufficient security levels may lead to security vulnerabilities. In low-risk transactions, setting the threshold too high may increase the difficulty and time cost of user verification, reducing overall system performance.
[0006] Furthermore, existing multilingual voiceprint authentication systems typically deploy independent models for each language. While this deployment approach achieves multilingual support to a certain extent, it results in significant resource waste. Each language model requires independent computing resources and storage space. As the number of languages increases, the hardware resources required for the system grow linearly, increasing operational costs for enterprises. Furthermore, the independence of each model and the lack of effective coordination mechanisms lead to high system response latency, making it impossible to meet the stringent real-time requirements of cross-border financial transactions.
[0007] In summary, existing voiceprint authentication systems have many flaws in language processing, security strategies, and resource utilization, making it difficult to meet the growing demand for cross-border financial transactions. Therefore, developing a new voiceprint authentication technology that can solve the above problems has important practical significance and application value. Summary of the Invention
[0008] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a cross-border financial transaction control method, device, computer equipment and storage medium.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] Cross-border financial transaction control methods include:
[0011] Obtain the content of the user's voice input to obtain input information;
[0012] Based on the multilingual voiceprint feature model, the input information is segmented into different language segments, and the different language segments are weighted and fused to form a voiceprint feature vector;
[0013] The voiceprint matching threshold is calculated by combining the voiceprint feature vector and scene perception information;
[0014] Determine whether the voiceprint matching threshold is greater than or equal to the set threshold;
[0015] If the voiceprint matching threshold is greater than or equal to the set threshold, the cross-border transaction is executed.
[0016] A further technical solution is: obtaining the content of the user's voice input to obtain input information includes:
[0017] The cross-border transaction content of the user's multi-language mixed voice is obtained through the terminal to obtain input information.
[0018] A further technical solution is as follows: the multilingual voiceprint feature model is used to segment the input information into different language segments, and the different language segments are weighted and fused to form a voiceprint feature vector, including:
[0019] Segment the input information into different language segments;
[0020] Input different language segments into the multilingual voiceprint feature model to obtain extracted features;
[0021] The extracted features are weighted and fused to form a voiceprint feature vector.
[0022] A further technical solution is: the multilingual voiceprint feature model includes a shared encoder and an adapter, the shared encoder is used to extract cross-language common features, and the adapter is used to capture language specificity.
[0023] A further technical solution is that the cross-language common features and the language-specific features are combined to form the extracted features.
[0024] A further technical solution is: the weighted fusion of the extracted features to form a voiceprint feature vector includes:
[0025] The duration of different language segments and the weighted fusion of extracted features are combined to form a voiceprint feature vector.
[0026] Its further technical solution is: the scene perception information includes: user nationality, transaction type and IP geographic location.
[0027] The present invention also provides a cross-border financial transaction control device, comprising:
[0028] An acquisition unit, configured to acquire the content of the user's voice input to obtain input information;
[0029] A segmentation and fusion unit is used to segment the input information into different language segments based on the multilingual voiceprint feature model, and to weightedly fuse the different language segments to form a voiceprint feature vector;
[0030] A combined calculation unit is used to calculate the voiceprint matching threshold by combining the voiceprint feature vector and scene perception information;
[0031] A judgment unit, used to judge whether the voiceprint matching threshold is greater than or equal to a set threshold;
[0032] The execution unit is used to execute the cross-border transaction if the voiceprint matching threshold is greater than or equal to the set threshold.
[0033] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0034] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0035] Compared with the existing technology, the beneficial effects of the present invention are as follows: by constructing a multilingual voiceprint feature model, the content of the user's voice input can be accurately segmented, different language segments can be effectively identified and weighted and fused to form a voiceprint feature vector, so that the system can support mixed input of multiple languages such as Chinese, English, Cantonese, Malay, etc. When facing a complex language environment, the system can accurately extract the user's voiceprint features, greatly improving the accuracy and success rate of verification; in addition, by combining the voiceprint feature vector and scene perception information to calculate the voiceprint matching threshold, a dynamic threshold mechanism is realized. Through comprehensive analysis of these factors, the system can accurately judge the security risk level of the current transaction scenario and respond accordingly. The voiceprint matching threshold is adjusted accordingly. In high-risk cross-border transfer transactions, the system will automatically increase the voiceprint matching threshold and increase the strictness of verification, thereby effectively reducing the false acceptance rate and preventing the occurrence of illegal transactions. In addition, by building a unified multilingual voiceprint feature model, centralized processing and analysis of voiceprint information in multiple languages is achieved, avoiding the waste of resources caused by the independent deployment of multiple language models. At the same time, the weighted fusion method enables the system to extract voiceprint features more efficiently, reducing unnecessary calculations, thereby improving the system's response speed. In cross-border financial transactions, a fast response speed can improve user experience, reduce transaction waiting time, and increase user trust in the system.
[0036] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 A schematic diagram of an application scenario of the cross-border financial transaction control method provided by an embodiment of the present invention;
[0039] Figure 2 A schematic diagram of a flow chart of a cross-border financial transaction control method provided by an embodiment of the present invention;
[0040] Figure 3 A schematic block diagram of a cross-border financial transaction control device provided by an embodiment of the present invention;
[0041] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0043] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0044] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0045] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0046] See also Figure 1 and Figure 2 , Figure 1 Schematic diagram of an application scenario of the cross-border financial transaction control method provided in an embodiment of the present invention. Figure 2This is a schematic flow chart of a cross-border financial transaction control method provided by an embodiment of the present invention. This cross-border financial transaction control method is applied to a server that interacts with a terminal for data exchange. By fully leveraging the parallel computing capabilities of multi-core processors, it accurately segments the content of user voice input, effectively identifies different language segments, and weightedly fuses them to form a voiceprint feature vector. This enables the system to support mixed input in multiple languages, including Chinese, English, Cantonese, and Malay. In complex language environments, the system can accurately extract the user's voiceprint features, significantly improving the accuracy and success rate of verification. In this method, first, by constructing a multilingual voiceprint feature model, the content of the user's voice input can be accurately segmented, and different language segments can be effectively identified and weighted and fused to form a voiceprint feature vector, so that the system can support mixed input of multiple languages such as Chinese, English, Cantonese, Malay, etc. When faced with complex language environments, the system can accurately extract the user's voiceprint features, greatly improving the accuracy and success rate of verification; in addition, by combining the voiceprint feature vector and scene perception information to calculate the voiceprint matching threshold, a dynamic threshold mechanism is realized. Through comprehensive analysis of these factors, the system can accurately judge the security risk level of the current transaction scenario and adjust the voiceprint matching threshold accordingly. In high-risk cross-border transfer transactions, the system will automatically increase the voiceprint matching threshold and increase the strictness of verification, thereby effectively reducing the false acceptance rate and preventing the occurrence of illegal transactions.
[0047] Figure 2 This is a flow chart of the cross-border financial transaction control method provided by an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S160.
[0048] S110: Acquire the content of the user's voice input to obtain input information;
[0049] Specifically, in cross-border financial transaction scenarios, highly sensitive microphone hardware devices can be integrated into various transaction terminals (such as mobile banking apps, smart counters, and self-service terminals). For example, mobile banking apps can utilize the phone's built-in microphone, and smart counters can be equipped with professional directional microphones to ensure clear and accurate capture of user voice input. Furthermore, the collection environment can be optimized, such as by installing soundproofing materials around smart counters to reduce external noise interference. In mobile banking apps, when users initiate voice input, they can be prompted to select a relatively quiet environment, and software algorithms can be used to perform preliminary noise reduction on the collected voice signals to remove interfering factors such as background noise and echoes. Furthermore, the collected voice signals are stored in a unified audio format, such as the common WAV or MP3 formats, for subsequent processing and analysis. At the same time, relevant information such as the time and location of voice collection is recorded as reference data for subsequent transaction scenario perception.
[0050] In other words, by using highly sensitive microphone hardware and an optimized acquisition environment, external noise interference can be effectively reduced, ensuring that the collected voice signals are clear and accurate. At the same time, the application of voice activity detection algorithms and feature extraction algorithms can accurately segment voice signals and extract effective voice features, avoiding subsequent processing errors caused by incomplete voice signals or inaccurate feature extraction. For example, in a noisy airport environment, the above-mentioned technical means can still accurately capture the user's voice input, providing reliable basic data for subsequent voiceprint authentication.
[0051] In one embodiment, obtaining the content of the user's voice input to obtain input information includes:
[0052] The cross-border transaction content of the user's multi-language mixed voice is obtained through the terminal to obtain input information.
[0053] Specifically, a dedicated voice acquisition module is developed within the terminal device's operating system or related applications. This module features real-time voice data acquisition, encoding, and transmission capabilities, dynamically adjusting the sampling rate and bit rate based on varying network environments and device performance to ensure voice data quality and transmission efficiency. For example, in the case of a weak network signal, the sampling rate is appropriately lowered to reduce data volume and ensure stable voice data transmission. Furthermore, acoustic and language models are constructed for multiple languages (such as Chinese, English, Cantonese, and Malay). The acoustic model describes the mapping relationship between voice signals and phonemes, while the language model performs semantic analysis of phoneme sequences. Trained with large amounts of multilingual voice data, the model accurately identifies the speech features and grammatical structures of different languages. The acoustic and language models are then used to read the user's multilingual mixed voice content for cross-border transactions to obtain input information.
[0054] In other words, by building acoustic and language models encompassing multiple languages, the system can accurately recognize and process mixed multilingual speech. Whether it's a mix of Chinese and English or Cantonese and Malay, the system effectively extracts speech features to support cross-border financial transactions. This enables banks and other financial institutions to better serve global customers and expand their international business. For example, in Southeast Asia, users can conduct cross-border transactions in multiple languages, and the system can accurately recognize and process them, improving transaction success rates and user experience.
[0055] S120, segmenting the input information into different language segments based on the multilingual voiceprint feature model, and weightedly fusing the different language segments to form a voiceprint feature vector;
[0056] Specifically, a large amount of speech data is collected, encompassing multiple languages (such as Chinese, English, Cantonese, and Malay), encompassing people of varying genders, ages, accents, and pronunciation habits. Each speech segment is annotated, clearly indicating the language type, start and end time points, and the corresponding speaker identity information. For example, for a speech segment containing a mixture of Chinese and English, such as "I want to transfer money (Chinese) to my account (English)", the specific time periods of the Chinese and English portions need to be accurately annotated. Deep learning algorithms, such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or a combination of convolutional neural networks (CNNs) and RNNs, are trained on the annotated multilingual speech data. The model aims to learn the differences and commonalities in voiceprint features across languages, as well as the characteristic patterns of the voiceprints of the same speaker across different languages. For example, through training on this large amount of data, the model can identify the tonal characteristics of Chinese and the pronunciation patterns of phonemes in English, and establish a correlation between the two.
[0057] Feature extraction is performed on the input multilingual speech mixture. Common features include Mel-Frequency Cepstral Coefficients (MFCCs) and Linear Prediction Coefficients (LPCs). These features effectively represent the spectral characteristics of the speech signal and provide a basis for subsequent language segmentation. For example, the MFCC feature extraction algorithm is used to convert the input speech signal into a sequence of MFCC feature vectors.
[0058] The extracted feature vector sequence is input into a trained multilingual voiceprint feature model. The model will segment the input information in real time based on the differences in voiceprint features between different languages. Specifically, the model calculates the similarity between the feature vector at each time point and the voiceprint feature templates of different languages. When the similarity exceeds a certain threshold, the time point is determined to belong to the corresponding language segment. For example, if the feature vector of a certain time period is detected to have a high similarity with the Chinese voiceprint feature template, the time period is marked as a Chinese segment.
[0059] In addition, each language segment is weighted based on factors such as its importance, duration, and clarity. For example, segments containing key information, such as the transaction amount, can be assigned a higher weight, while segments containing modal particles or redundant information can be assigned a lower weight. Weights can be determined through statistical analysis or machine learning algorithms. For example, by analyzing large amounts of cross-border transaction voice data, the frequency and importance of different types of language segments in transactions can be statistically analyzed to determine the corresponding weights. The feature vectors of the different language segments are then weighted and fused according to the determined weights. This can be achieved by using simple weighted averaging or more complex fusion algorithms, such as principal component analysis (PCA) or linear discriminant analysis (LDA). For example, for two language segments A and B, whose feature vectors are vA and vB, and whose weights are wA and wB, respectively, the fused voiceprint feature vector v can be expressed as v = wAvA + wBvB.
[0060] In other words, after being trained on a large amount of multilingual speech data, the multilingual voiceprint feature model can accurately identify switching points and boundaries between different languages. For example, when processing mixed Chinese and English speech, it can accurately segment the Chinese and English portions, avoiding subsequent processing errors caused by inaccurate language boundary identification. This technology is adaptable to language environments with varying accents, speaking speeds, and pronunciation habits. The model can effectively segment both Chinese with a regional accent and English spoken quickly. For example, when a user with a Cantonese accent says "I want to go to Hong Kong (Cantonese) to Hong Kong (English)", the system can accurately segment the Cantonese and English segments. Furthermore, by weightedly fusing different language segments to form a voiceprint feature vector, it is possible to comprehensively utilize voiceprint information from multiple languages and enrich the representation capabilities of the voiceprint feature. For example, in cross-border transaction scenarios, users' speech may contain both Chinese and English. The fused voiceprint feature vector can more comprehensively reflect the user's voiceprint characteristics, improving the accuracy of voiceprint matching. The weighted fusion of different language segments can reduce the influence of single language features on voiceprint recognition and improve the system's robustness to speaker variations. For example, when users use different pronunciation habits in different language environments, the fused voiceprint feature vector can better adapt to these changes and ensure the stability of voiceprint recognition.
[0061] In one embodiment, the multilingual voiceprint feature model is used to segment the input information into different language segments, and weightedly fuse the different language segments to form a voiceprint feature vector, including:
[0062] Segment the input information into different language segments;
[0063] Input different language segments into the multilingual voiceprint feature model to obtain extracted features;
[0064] The extracted features are weighted and fused to form a voiceprint feature vector.
[0065] Specifically, a series of rules are developed for language segmentation based on the pronunciation rules and vocabulary characteristics of different languages. For example, for mixed Chinese and English speech, we can take advantage of the fact that Chinese words are generally composed of multiple syllables, while English words are composed of letter combinations. The language segments can be segmented by detecting the switching between syllables and letter combinations.
[0066] The segmented language segments are fed into a trained multilingual voiceprint feature model. The model processes each segment and extracts a feature vector that characterizes the speaker's voiceprint. For example, for a Chinese speech segment, the model can extract the speaker's voiceprint feature vector in a Chinese context. This vector contains information such as the speaker's timbre, pitch, and speaking speed.
[0067] The weight of each language segment's feature vector is determined based on its importance in the input information. For example, in a cross-border transaction scenario, if key information such as the transaction amount is expressed in a certain language, the feature vector weight of that language segment can be set higher. The voice quality of different language segments, such as signal-to-noise ratio and clarity, is also considered. The feature vector weight of language segments with better voice quality can be increased accordingly to ensure the accuracy of the fused voiceprint feature vector.
[0068] The feature vectors of different language segments are weighted averaged according to the determined weights to obtain the fused voiceprint feature vector.
[0069] In other words, by combining model-based and rule-based language segmentation methods, it is possible to adapt to complex language environments with varying accents, speaking speeds, and levels of language mixing. For example, when processing multilingual mixed speech with regional accents, the system can still accurately segment different language segments. Compared to traditional single-language segmentation methods, this method can more comprehensively consider the characteristics of different languages, reducing segmentation errors caused by blurred language boundaries or similar language features. In addition, weighted fusion of feature vectors of different language segments can comprehensively utilize voiceprint information from multiple languages and enrich the representation capabilities of voiceprint features. For example, in cross-border transactions, a user's voice may contain both Chinese and English. The fused voiceprint feature vector can more comprehensively reflect the user's voiceprint characteristics and improve the accuracy of voiceprint matching.
[0070] In one embodiment, the multilingual voiceprint feature model includes a shared encoder and an adapter, wherein the shared encoder is used to extract cross-language common features, and the adapter is used to capture language specificity.
[0071] Specifically, the multilingual voiceprint feature fusion model includes a cross-language transfer learning architecture: a shared encoder and adapter are designed. The model structure is as follows:
[0072] h shared =Encoder(X audio ),h lang =Adapter k (h shared );
[0073] Among them, Adapter k Corresponding to the kth language (such as English and Cantonese), the shared encoder is used to extract cross-language common features, and the adapter is used to capture language specificity.
[0074] Parameter Description:
[0075] X audio : The original input speech signal (such as waveform or spectrogram).
[0076] Encoder: A shared encoder used to extract common voiceprint features across languages (such as tone and pronunciation habits).
[0077] h shared : Share the cross-lingual feature vector output by the encoder.
[0078] Adapter k : An adapter for the kth language (e.g., English, Cantonese) to capture language-specific features (e.g., English linking, Cantonese tones).
[0079] h lang : Language-related voiceprint features output by the adapter.
[0080] In other words, the shared encoder can extract common features across languages. These features are not restricted to a specific language and can reflect the essential voiceprint characteristics of the speaker. For example, speech signals in different languages may share commonalities in pitch and timbre. The shared encoder can capture these commonalities, thereby improving the accuracy of voiceprint recognition. Furthermore, the adapter can capture language-specific features that can offset the shortcomings of the shared encoder in handling specific languages. For example, some languages may have unique pronunciation rules or intonation characteristics. The adapter can adjust and transform these specific features, making the model more adaptable to speech signals in different languages and further improving the accuracy of voiceprint recognition. Furthermore, because the shared encoder extracts common features across languages, when encountering a new language, only a corresponding adapter needs to be added, without retraining the entire model. This gives the model strong generalization capabilities and allows it to quickly adapt to new languages. In real applications, speech signals may be affected by various factors, such as speaker accent variations and environmental noise. The combination of the shared encoder and the adapter enables the model to better cope with these changes, improving its robustness and generalization capabilities.
[0081] In one embodiment, the weighted fusion of the extracted features to form the voiceprint feature vector includes: combining the duration of different language segments with the weighted fusion of the extracted features to form the voiceprint feature vector. The cross-language common features and the language-specific features are combined to form the extracted features.
[0082] Specifically, a language boundary detection model based on deep learning, such as a recurrent neural network (RNN) or a convolutional neural network (CNN) combined with a long short-term memory network (LSTM), is used to detect language boundaries for mixed Chinese and English speech. The model input is the Mel-frequency cepstral coefficient (MFCC) feature of the speech signal, and the output is the language label (Chinese or English) for each time frame. According to the language boundary detection results, the mixed Chinese and English speech is segmented into Chinese segments and English segments. Specific feature extraction adapters are designed for Chinese and English respectively. These adapters can be lightweight neural network modules, such as fully connected layers or convolutional layers, which are used to extract features from the segmented speech segments. The segmented Chinese and English segments are respectively input into the corresponding adapters to extract language-specific feature vectors. The voiceprint feature vector is calculated specifically by the following formula:
[0083] v fusion =αv zh +(1-α)v en ,
[0084] Among them, T zh and T en The duration ratios of Chinese and English segments respectively.
[0085] Parameter Description:
[0086] T zh : Duration of the Chinese segment in the speech (unit: seconds).
[0087] T en : Duration of English segments in audio.
[0088] α: The weight of the Chinese duration ratio, used to dynamically adjust the fusion ratio of Chinese and English features.
[0089] v zh : Feature vectors extracted from Chinese speech segments using the Chinese adapter.
[0090] v en : Feature vectors extracted from English speech segments through the English adapter.
[0091] v fusion : The fused multilingual voiceprint feature vector.
[0092] That is, the h extracted by the shared encoder shared is the input of the adapter, and the output of different language adapters v zh and v en The final voiceprint feature vector is generated through a fusion formula. This involves extracting common features across different languages to reduce recognition errors caused by language differences. Language-specific features are then combined to more accurately identify voiceprint features in different languages, improving recognition accuracy. Furthermore, by considering the duration of different language segments and assigning appropriate weights, recognition bias caused by duration differences is reduced. The fusion of cross-lingual common features and language-specific features enhances the robustness of the voiceprint feature vector, ensuring stable operation in various language environments.
[0093] S130, calculating a voiceprint matching threshold by combining the voiceprint feature vector and scene perception information;
[0094] Specifically, the scenario perception information includes: user nationality, transaction type and IP geographic location. More specifically, when a user registers, the user is required to fill in the nationality information and store it in the database for subsequent calls. Or by calling a third-party API (such as a geocoding service), the user's nationality is inferred based on the user's IP address or device information as supplementary information. In addition, the transaction type of the current transaction (such as transfer, consumption, recharge, etc.) is obtained from the transaction system, and the transaction type is determined according to the specific circumstances of the transaction. Or during the transaction process, the user is required to select the transaction type to ensure the accuracy of the transaction type. In addition, based on the user's IP address, geographic location information is obtained through a third-party service to determine the user's approximate location. Or in combination with the geographic information database, the obtained IP geographic location is further verified and refined to improve the accuracy of the geographic location.
[0095] In addition, based on historical data and experience, a static voiceprint matching threshold (such as 0.7) is set as the basic threshold to provide a reference for subsequent dynamic adjustments. The matching threshold is then adjusted based on the user's nationality. For example, for users from high-risk countries or regions, the threshold can be increased (such as 0.8) to improve the strictness of the match. The matching threshold is then adjusted based on the transaction type. For example, for large transactions or high-risk transaction types, the threshold can be increased to ensure the security of the transaction. The matching threshold can also be adjusted based on the user's IP geographic location. For example, for IP addresses that are significantly different from the user's usual geographic location, the threshold can be increased to reduce the risk of misidentification. The final voiceprint matching threshold is obtained by taking a weighted average of the basic threshold and the threshold adjusted based on scene perception information. For example, the weight of the basic threshold can be set to 0.6, and the weight of the threshold adjusted based on scene perception information can be set to 0.4, and dynamic adjustments can be made based on specific scenarios.
[0096] In other words, by combining context-aware information such as user nationality, transaction type, and IP location, a more comprehensive assessment of transaction risk can be made, allowing the voiceprint matching threshold to be adjusted to improve matching accuracy and reliability. Furthermore, by dynamically adjusting the matching threshold based on different context-aware information, the system can adapt to different transaction environments and reduce the risk of misidentification and missed identification. For users from high-risk countries or regions, large transactions, or high-risk transaction types, the matching threshold is increased to improve transaction security and prevent fraud. Furthermore, by verifying the user's IP location, the risk of fraudulent transactions is reduced. For example, if a user typically conducts transactions in Beijing and suddenly encounters a transaction from another region, the system can increase the matching threshold and require the user to undergo additional identity verification.
[0097] In one embodiment, the optimal voiceprint model and matching threshold are selected in real time based on the user's nationality, transaction type (e.g., remittance, inquiry), and IP location:
[0098] Threshold=β·RiskScore+(1-β)·LatencyConstraint;
[0099] Among them, RiskScore∈[0,1]: Risk score based on transaction amount and destination (the threshold is increased by 20% in high-risk scenarios). LatencyConstraint: In low-latency demand scenarios (such as real-time payments), the threshold is moderately relaxed to ensure speed.
[0100] Parameter Description:
[0101] β: A trade-off factor (value range [0, 1]) used to balance risk and latency requirements (e.g., β = 0.8 for high-risk scenarios and β = 0.3 for low-latency scenarios).
[0102] RiskScore: Transaction risk score (based on amount, destination, etc.), value range [0,1].
[0103] LatencyConstraint: Latency constraint coefficient (if low latency is required in real-time payment scenarios, the threshold can be relaxed).
[0104] Threshold: The final dynamically adjusted voiceprint matching threshold (the higher the threshold, the stricter the verification).
[0105] Specifically, the RiskScore and LatencyConstraint in the dynamic threshold formula affect the strictness of the match, which indirectly depends on the accuracy of the fused features. By pre-installing a lightweight multilingual voiceprint feature fusion model on the user-end device (such as a mobile phone), only the encrypted voiceprint features are uploaded to the cloud, reducing latency to <200ms.
[0106] S140, determining whether the voiceprint matching threshold is greater than or equal to a set threshold;
[0107] The voiceprint matching threshold is used to match the user's voiceprint template, and the dynamic threshold (i.e., the voiceprint matching threshold) needs to be compared with the set threshold to determine whether the verification is passed.
[0108] Specifically, according to the system's security policies and business needs, corresponding thresholds are set for transactions of different risk levels. For example, for ordinary transactions, the threshold is set at 0.75; for high-risk transactions, the threshold is set at 0.85. In addition, clear rules are formulated to specify how to select and apply the thresholds in different scenarios. For example, when the transaction amount exceeds a certain amount, the threshold for high-risk transactions is automatically selected. The calculated voiceprint matching threshold (dynamic threshold) is compared with the pre-set threshold in real time. If the voiceprint matching threshold is greater than or equal to the set threshold, the voiceprint match is considered successful and the user passes the verification; otherwise, the voiceprint match is considered failed and the user fails the verification.
[0109] In other words, dynamically adjusting the voiceprint matching threshold based on different scenario-based information can more accurately assess the authenticity of the user's identity. For example, for high-risk transactions, increasing the threshold can increase the stringency of verification and prevent fraud. By using reasonable threshold comparisons, false and missed identifications caused by similar voiceprint features or environmental factors can be reduced, thereby improving the reliability of identity verification.
[0110] S150: If the voiceprint matching threshold is greater than or equal to the set threshold, the cross-border transaction is executed;
[0111] Specifically, if the voiceprint matching threshold is greater than or equal to the set threshold, the verification is considered successful and the system executes the cross-border transaction. For example, the transaction processing module is called to complete operations such as fund transfer and transaction record storage.
[0112] In other words, accurate voiceprint matching and threshold comparison can reduce the cost of manual review and improve the system's automation. Through risk assessment and dynamic threshold adjustment, cross-border transaction risks can be more effectively managed, reducing losses caused by fraudulent transactions and thus reducing operating costs.
[0113] S160: If the voiceprint matching threshold is less than the set threshold, manual review is triggered.
[0114] Specifically, if the voiceprint matching threshold is lower than a set threshold, the verification is deemed risky and a manual review process is triggered. The system then assigns the review task to an appropriate human reviewer. This intelligent assignment is based on factors such as the reviewer's expertise and workload, ensuring timely processing of review tasks. Furthermore, the reviewer is provided with relevant audit information, including the user's transaction details, voiceprint matching score, and voice recordings, enabling a comprehensive assessment.
[0115] That is to say, by comparing the voiceprint matching threshold with the set threshold, transactions with risks in identity verification can be detected in a timely manner, triggering manual review and effectively preventing security threats such as impersonation and fraud. For example, when criminals attempt to use forged voiceprints for cross-border transactions, due to the low voiceprint matching score, the system will trigger manual review to prevent the transaction from proceeding. In addition, by combining voiceprint technology and manual review, a multi-level security verification mechanism is achieved. Voiceprint technology can quickly screen out most normal transactions, while for transactions with risks, manual review can provide more in-depth analysis and judgment to further improve the security of transactions.
[0116] To facilitate the calculation of this technical solution, the following examples are provided:
[0117] Suppose a Hong Kong user conducts a cross-border transfer of mixed Chinese and English voice through mobile banking.
[0118] User voice input: "Transfer¥50000to John’s UK account, 转账五万元到约翰的英国账户".
[0119] Then perform voice segmentation and processing:
[0120] Segmented into an English segment ("Transfer...account") and a Chinese segment ("转账...账户"), with a duration ratio α = 0.6 (Chinese accounts for 60%).
[0121] Adjustment of the voiceprint matching threshold:
[0122] Since the transaction amount is large (¥50000) and the destination is a high-risk country (UK), RiskScore = 0.9, β = 0.7, the calculated threshold is increased to 0.85.
[0123] Result verification:
[0124] The system conducts a risk assessment based on the transaction amount (50000) and the country of the transaction counterpart (high-risk country), determines that this transaction is a high-risk transaction, and adjusts the set threshold to 0.8 according to the pre-set risk level and set threshold mapping table. Among them, since the voiceprint matching threshold of 0.85 exceeds the set threshold of 0.8, the verification passes and the transaction is completed, resulting in an end-to-end delay of only 180 ms.
[0125] The above-mentioned cross-border financial transaction control method, by constructing a multilingual voiceprint feature model, can accurately segment the content of user voice input, effectively identify different language segments, and weightedly fuse them to form a voiceprint feature vector. This enables the system to support mixed input in multiple languages, such as Chinese, English, Cantonese, and Malay. In complex language environments, the system can accurately extract the user's voiceprint features, greatly improving the accuracy and success rate of verification. For example, in cross-border transactions involving Southeast Asia, users may communicate in both Chinese and Malay. This technology can accurately identify and process voiceprint information in both languages, ensuring smooth transactions and effectively addressing the shortcomings of existing technologies in terms of multilingual compatibility.
[0126] In addition, by combining voiceprint feature vectors and scene perception information to calculate the voiceprint matching threshold, a dynamic threshold mechanism is implemented. Through comprehensive analysis of these factors, the system can accurately determine the security risk level of the current transaction scenario and adjust the voiceprint matching threshold accordingly. In high-risk cross-border transfer transactions, the system automatically increases the voiceprint matching threshold and increases the strictness of verification, thereby effectively reducing the false acceptance rate and preventing illegal transactions. For example, when a user makes a large cross-border transfer, the system will adjust the voiceprint matching threshold to a higher level based on information such as the transaction amount and transaction time, ensuring that only strictly verified users can complete the transaction, greatly improving transaction security.
[0127] In addition, by building a unified multilingual voiceprint feature model, centralized processing and analysis of voiceprint information in multiple languages is achieved, avoiding the resource waste caused by independently deploying multiple language models. At the same time, the weighted fusion method enables the system to extract voiceprint features more efficiently, reducing unnecessary computation and thus improving the system's response speed. In cross-border financial transactions, fast response speed can improve user experience, reduce transaction waiting time, and increase user trust in the system. For example, when a user makes an urgent cross-border payment, the system can quickly and accurately complete voiceprint verification, promptly process the transaction request, and ensure that the funds arrive in time.
[0128] Furthermore, this technology's multilingual compatibility and dynamic threshold mechanism provide strong support for banks to rapidly expand their cross-border business. With the continuous development of cross-border financial markets, banks need to continuously expand their business scope and attract more international customers. This technology can meet the multilingual needs of users in different countries and regions, allowing banks to easily provide cross-border financial services in multiple languages. Furthermore, the dynamic threshold mechanism can be flexibly adjusted based on different business scenarios and risk levels to ensure transaction security and compliance. Banks no longer need to develop separate voiceprint authentication systems for transactions in different languages and risk levels, significantly reducing the cost and difficulty of business expansion and improving its efficiency. For example, if a bank plans to expand into the Southeast Asian market, this technology can help the bank quickly establish a voiceprint authentication system that supports multiple Southeast Asian languages, providing local users with convenient and secure cross-border financial services and promoting the rapid development of the bank's cross-border business.
[0129] Figure 3 : is a schematic block diagram of a cross-border financial transaction control device 300 provided by an embodiment of the present invention. Figure 3 As shown, corresponding to the above cross-border financial transaction control method, the present invention also provides a cross-border financial transaction control device 300. The cross-border financial transaction control device 300 includes a unit for executing the above cross-border financial transaction control method, and the device can be configured in a server. Specifically, please refer to Figure 3 The cross-border financial transaction control device 300 includes an acquisition unit 301, a segmentation and fusion unit 302, a combined calculation unit 303, a judgment unit 304 and an execution unit 305.
[0130] The acquisition unit 301 is used to acquire the content of the user's voice input to obtain input information;
[0131] The segmentation and fusion unit 302 is configured to segment the input information into different language segments based on the multilingual voiceprint feature model, and perform weighted fusion of the different language segments to form a voiceprint feature vector;
[0132] The combined calculation unit 303 is used to calculate the voiceprint matching threshold by combining the voiceprint feature vector and the scene perception information;
[0133] The judging unit 304 is used to judge whether the voiceprint matching threshold is greater than or equal to a set threshold;
[0134] The execution unit 305 is configured to execute the cross-border transaction if the voiceprint matching threshold is greater than or equal to a set threshold.
[0135] In one embodiment, the acquisition unit 301 acquires the cross-border transaction content in the multi-language mixed voice of the user through the terminal to obtain input information.
[0136] In one embodiment, the segmentation and fusion unit 302 includes:
[0137] A segmentation subunit, used to segment the input information into different language segments;
[0138] An input subunit, used to input different language segments into the multilingual voiceprint feature model to obtain extracted features;
[0139] The fusion subunit is used to perform weighted fusion of the extracted features to form a voiceprint feature vector.
[0140] In one embodiment, the multilingual voiceprint feature model includes a shared encoder and an adapter, wherein the shared encoder is used to extract cross-language common features, and the adapter is used to capture language specificity.
[0141] In one embodiment, the cross-language common features and the language-specific features are combined to form the extracted features.
[0142] In one embodiment, the fusion subunit forms a voiceprint feature vector by combining the duration of different language segments and extracting weighted features.
[0143] In one embodiment, the scenario-aware information includes: user nationality, transaction type, and IP geographic location.
[0144] In one embodiment, the apparatus further includes: a triggering unit 306, configured to trigger manual review if the voiceprint matching threshold is less than a set threshold.
[0145] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned cross-border financial transaction control device 300 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of the description, it will not be repeated here.
[0146] The cross-border financial transaction control device 300 can be implemented in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer device shown.
[0147] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.
[0148] See Figure 4 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0149] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can cause the processor 502 to execute a cross-border financial transaction control method.
[0150] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0151] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a cross-border financial transaction control method.
[0152] The network interface 505 is used to communicate with other devices through the network. Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0153] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:
[0154] Obtain the content of the user's voice input to obtain input information; segment the input information into different language segments based on the multilingual voiceprint feature model, and weightedly fuse the different language segments to form a voiceprint feature vector; calculate the voiceprint matching threshold by combining the voiceprint feature vector and scene perception information; determine whether the voiceprint matching threshold is greater than or equal to the set threshold; if the voiceprint matching threshold is greater than or equal to the set threshold, execute the cross-border transaction.
[0155] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0156] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0157] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:
[0158] Obtain the content of the user's voice input to obtain input information; segment the input information into different language segments based on the multilingual voiceprint feature model, and weightedly fuse the different language segments to form a voiceprint feature vector; calculate the voiceprint matching threshold by combining the voiceprint feature vector and scene perception information; determine whether the voiceprint matching threshold is greater than or equal to the set threshold; if the voiceprint matching threshold is greater than or equal to the set threshold, execute the cross-border transaction.
[0159] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0160] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0161] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0162] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0163] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0164] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A cross-border financial transaction control method, characterized in that: include: Obtain the content of the user's voice input to obtain input information; Based on the multilingual voiceprint feature model, the input information is segmented into different language segments, and the different language segments are weighted and fused to form a voiceprint feature vector; The voiceprint matching threshold is calculated by combining the voiceprint feature vector and scene perception information; Determine whether the voiceprint matching threshold is greater than or equal to the set threshold; If the voiceprint matching threshold is greater than or equal to the set threshold, the cross-border transaction is executed.
2. The cross-border financial transaction control method according to claim 1, characterized in that: The acquiring of the content of the user's voice input to obtain input information includes: The cross-border transaction content of the user's multi-language mixed voice is obtained through the terminal to obtain input information.
3. The cross-border financial transaction control method according to claim 1, characterized in that: The multilingual voiceprint feature model is used to segment the input information into different language segments, and weightedly fuse the different language segments to form a voiceprint feature vector, including: Segment the input information into different language segments; Input different language segments into the multilingual voiceprint feature model to obtain extracted features; The extracted features are weighted and fused to form a voiceprint feature vector.
4. The cross-border financial transaction control method according to claim 3, characterized in that: The multilingual voiceprint feature model includes a shared encoder and an adapter. The shared encoder is used to extract cross-language common features, and the adapter is used to capture language specificity.
5. The cross-border financial transaction control method according to claim 4, characterized in that: The cross-language common features and the language-specific features are combined to form the extracted features.
6. The cross-border financial transaction control method according to claim 5, characterized in that: The step of weighting and fusing the extracted features to form a voiceprint feature vector includes: The duration of different language segments and the weighted fusion of extracted features are combined to form a voiceprint feature vector.
7. The cross-border financial transaction control method according to claim 1, characterized in that: The scenario perception information includes: user nationality, transaction type and IP geographic location.
8. A cross-border financial transaction control device, characterized in that: include: An acquisition unit, configured to acquire the content of the user's voice input to obtain input information; A segmentation and fusion unit is used to segment the input information into different language segments based on the multilingual voiceprint feature model, and to weightedly fuse the different language segments to form a voiceprint feature vector; A combined calculation unit is used to calculate the voiceprint matching threshold by combining the voiceprint feature vector and scene perception information; A judgment unit, used to judge whether the voiceprint matching threshold is greater than or equal to a set threshold; The execution unit is used to execute the cross-border transaction if the voiceprint matching threshold is greater than or equal to the set threshold.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.