Personal voice intelligent terminal system for fund transaction room

Through real-time voice acquisition and adaptive semantic segmentation technology, the priority of speech analysis is dynamically adjusted, which solves the problem that the speech recognition system in the existing technology cannot adapt to the high-frequency trading environment, and realizes the accurate and timely execution of trading instructions, improving transaction efficiency and decision-making accuracy.

CN120279914AInactive Publication Date: 2025-07-08SHANGHAI SHOUEE TECH CO LTD

Patent Information

Application Number
CN202510782108.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing voice recognition system cannot adaptively adjust in high-frequency trading environments such as the capital trading room, resulting in the failure of voice command analysis or some content not being recognized, resulting in transaction errors.

Method used

Through real-time voice acquisition, preprocessing, speech feature analysis and urgency evaluation, adaptive semantic segmentation and priority control modules, the speech analysis priority is dynamically adjusted to ensure that trader instructions are executed accurately and promptly.

Benefits of technology

It improves trading efficiency and market response capabilities, avoids omissions or missed execution of instructions, and improves traders' operational efficiency and decision-making accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279914A_ABST
    Figure CN120279914A_ABST
Patent Text Reader

Abstract

The invention discloses a personal voice intelligent terminal system for a fund transaction room, which relates to the technical field of intelligent voice, and comprises a voice acquisition module, a voice preprocessing and structuring module, a voice feature analysis and emergency degree evaluation module, an intelligent recognition evaluation module and a self-adaptive semantic segmentation and priority regulation and control module, each voice instruction of a trader is quickly captured through a real-time voice acquisition system. A trader instruction is quickly captured through a real-time voice acquisition system, and the trader instruction is preprocessed and then converted into a structured data set. The voice emergency degree is quantified by extracting and analyzing key features, and the emergency of the instruction is intelligently evaluated through a machine learning model. After the system recognizes a quick and emergency instruction, the analysis priority of each section of voice is dynamically adjusted by utilizing a self-adaptive segmentation technology and context analysis, so that the trader instruction is ensured to be accurately and timely executed, the instruction is prevented from being omitted or mistakenly executed, and the transaction efficiency and the market response capability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice technology, and particularly to a personal voice intelligent terminal system for a fund trading room. Background Art

[0002] A personal voice intelligent terminal for a fund trading room refers to a terminal device integrating voice recognition, natural language processing, and intelligent control technologies, which is designed specifically for a fund trading room (such as a trading environment in financial markets like securities, foreign exchange, futures, etc.). This terminal uses voice input as the main way to interact with the system, enabling traders to perform tasks such as data query, trading operations, information reminder, and real-time monitoring through voice commands. Its core functions include the application of voice recognition technology to convert the trader's voice into executable instructions; natural language processing technology to understand and process complex voice inputs and conduct intention analysis; and an intelligent decision-making and feedback mechanism to assist traders in making efficient decisions in a busy trading environment.

[0003] Through such a voice intelligent terminal, traders can not only save time and reduce operation errors based on traditional manual input, but also achieve more efficient multitasking, improving work efficiency and the response speed of trading decisions. In a high-pressure and fast-response environment like a fund trading room, the role of the voice intelligent terminal is particularly important. It can quickly execute trades, query market quotes, set risk control alerts, etc. through voice commands, thereby enhancing the operation flexibility and decision-making accuracy of traders, and ultimately improving the overall trading efficiency and response speed.

[0004] The existing technology has the following deficiencies: In the existing voice recognition technology, especially in high-frequency trading environments such as fund trading rooms, the voice recognition system may have problems with non-self-adaptive adjustment when processing continuous voice inputs. When a trader quickly issues multiple voice commands in an emergency, the voice recognition system of the existing technology usually cannot adaptively adjust the recognition strategy according to the speed and complexity of the voice input. This deficiency in the lack of a dynamic adjustment mechanism causes the system to be unable to adapt to changes in the trader's voice input in real time, which may lead to the loss of some audio signals or signal interruptions, resulting in the failure of voice command parsing or partial content not being recognized.

[0005] For example, when a trader quickly issues an instruction like "Buy 1,000 shares of Company A's stock and sell 500 shares of Company B's stock", the existing system may not be able to perform self-adaptive processing according to the rapid changes in the voice input, resulting in the system failing to capture all the contents of the instruction. As a result, the system may only recognize "Buy 1,000 shares of Company A's stock" and miss the part of "Sell 500 shares of Company B's stock", causing incomplete trading instructions and further leading to trading errors.

[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] The object of the present invention is to provide a personal voice intelligent terminal system for a fund trading room, which quickly captures traders' instructions through a real-time voice acquisition system and converts them into a structured data set after preprocessing. By extracting and analyzing key features, quantifying the urgency of the voice, and intelligently evaluating the urgency of the instructions through a machine learning model. After the system identifies fast and urgent instructions, it uses adaptive segmentation technology and context analysis to dynamically adjust the parsing priority of each segment of voice, so as to ensure the accurate and timely execution of traders' instructions, avoid omission or misexecution of instructions, and improve trading efficiency and market response ability, so as to solve the problems in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solutions: A personal voice intelligent terminal system for a fund trading room, including a voice acquisition module, a voice preprocessing and structuring module, a voice feature analysis and urgency evaluation module, an intelligent recognition and evaluation module, and an adaptive semantic segmentation and priority regulation module: The voice acquisition module quickly captures each voice instruction of the trader through a real-time voice acquisition system; The voice preprocessing and structuring module preprocesses the collected voice signal and converts the preprocessed voice information into a structured data set; The voice feature analysis and urgency evaluation module extracts key features reflecting the fast and urgent input instructions of the trader from the data set, comprehensively analyzes the extracted key features, and quantifies the urgency of the voice; The intelligent recognition and evaluation module inputs the analyzed key features into a pre-trained machine learning model, and uses the model to intelligently evaluate the trader's instructions, so as to determine whether the instructions have the characteristics of being fast and urgent; The adaptive semantic segmentation and priority regulation module, when identifying that the instructions issued by the trader are fast and urgent, divides the continuous voice input of the trader into multiple paragraphs, and dynamically determines the parsing priority of each paragraph of voice in combination with the pauses, accents and context information of the current voice.

[0009] Preferably, the specific steps of preprocessing the collected voice signal and converting it into a structured data set are as follows: First, noise suppression and echo cancellation technologies are used to remove background noise and reflected sound to ensure the clarity of the voice signal; Next, the voice signal is enhanced and normalized to improve its volume and quality, making it more accurate in subsequent analysis; Subsequently, audio features of the speech are extracted through a feature extraction algorithm, the speech signal is segmented into multiple time frames, further divided into logical paragraphs, and pauses and tone changes are analyzed; Finally, the extracted speech features and information are converted into a structured data set and labeled, providing an accurate data basis for subsequent speech recognition, instruction parsing, and intelligent decision-making.

[0010] Preferably, key features reflecting the rapidity and urgency of the trader's input instructions are extracted from the data set. The extracted features include the comprehensive change in the pause frequency and pause duration in the speech and the fluctuation amplitude of the speech volume. The comprehensive change in the pause frequency and pause duration in the speech and the fluctuation amplitude of the speech volume are comprehensively analyzed under a detection window to generate a pause density reference value and a speech volume fluctuation reference value respectively, and the speech urgency is quantified through the pause density reference value and the speech volume fluctuation reference value.

[0011] Preferably, the specific steps for comprehensively analyzing the comprehensive change in the pause frequency and pause duration in the speech under a detection window to generate a pause density reference value are as follows: Within the detection window, all pause events in the speech are identified and recorded. Each pause event consists of the speech duration before the pause and the pause duration, forming a pause event sequence. The transition intensity factor of each pause event is calculated to measure the difference between the speech pause interval and the change in the speech duration. The calculation expression of the transition intensity factor is as follows: , where is the transition intensity factor, is the speech duration before the th pause, is the th pause duration, is a small constant, is the tension amplification factor, is the hyperbolic tangent function; Based on all the obtained transition intensity factors , a pause density reference value is generated to quantify the urgency of the speech instruction. The generation formula of the pause density reference value is as follows: , where is the pause density reference value, is the natural base, is the total number of effective pauses.

[0012] Preferably, the specific steps for comprehensively analyzing the fluctuation amplitude of the speech volume under a detection window to generate a speech volume fluctuation reference value are as follows: Within the detection window, the energy value of the trader's speech input signal is obtained and an input signal energy sequence is established , , where is the energy value of the -th frame of the voice signal, is the total number of voice frames. For the energy difference between adjacent frames, the local mutation amplitude is defined. The calculation expression of the local mutation amplitude is as follows: , where in the formula, is the local mutation amplitude, is the energy value of the -th frame of the voice signal, is the fluctuation sensitivity amplification factor; Based on the obtained local mutation amplitude , a voice volume fluctuation reference value is further constructed to comprehensively measure the volume fluctuation intensity and rhythm density of the voice signal within the detection window. The calculation formula of the voice volume fluctuation reference value is as follows: , where in the formula, is the voice volume fluctuation reference value, is the rhythm enhancement coefficient, is the continuous fluctuation times function.

[0013] Preferably, the analyzed pause density reference value and the voice volume fluctuation reference value are input into a pre-trained machine learning model, and an instruction urgency coefficient is generated through the machine learning model. The trader's instruction is intelligently evaluated through the instruction urgency coefficient, so as to judge whether the instruction has the characteristics of being fast and urgent.

[0014] Preferably, the instruction urgency coefficient generated when the trader's instruction is intelligently evaluated through a pre-trained machine learning model is compared and analyzed with a pre-set instruction urgency coefficient reference threshold to judge whether the instruction has the characteristics of being fast and urgent. The judgment logic is as follows: If the instruction urgency coefficient is greater than the pre-set instruction urgency coefficient reference threshold, it is judged that the instruction has the characteristics of being fast and urgent, indicating that the instruction issued by the trader is fast and urgent; If the instruction urgency coefficient is less than or equal to the pre-set instruction urgency coefficient reference threshold, it is judged that the instruction does not have the characteristics of being fast and urgent, indicating that the instruction issued by the trader does not have urgency.

[0015] Preferably, when it is recognized that the instruction issued by the trader is fast and urgent, the continuous voice input of the trader is divided into multiple paragraphs, and the specific steps of dynamically determining the parsing priority of each paragraph of voice in combination with the pauses, accents and context information of the current voice are as follows: When it is recognized that the instructions issued by the trader are fast and urgent, by analyzing the pauses, accents, and context information in the voice, the voice input of the trader is dynamically divided into multiple voice segments, and the parsing priority of each segment is adjusted according to the voice characteristics. The priority adjustment formula is as follows: , where is the priority score assigned to each voice segment, is the pause duration of the th voice segment in the voice input, is the intensity of the accented part of the th voice segment, is the degree of correlation between the context information of the th voice segment and the current trading environment, is the weight coefficient of the pause time, is the weight coefficient of the accent intensity, is the weight coefficient of the context relevance; After adjusting the priority of each voice segment based on the voice characteristics, the priorities of all segments are sorted to ensure that the instructions are executed according to the preset priority, so as to ensure that the urgent instructions can be processed first. The priority sorting formula is as follows: , where is the sorting position of the th voice segment in the final execution, is the priority of the th voice segment, is the total number of voice segments.

[0016] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: The present invention captures and preprocesses the voice signal in real time, extracts and quantifies the urgency of the trader's voice instructions, and determines the priority of the instructions through an intelligent evaluation model, so as to ensure that the urgent instructions can be preferentially recognized and processed. Further, through the adaptive segmentation technology and context analysis, the system can dynamically adjust the priority of voice parsing according to the pause, accent, and other information of the trader's voice, effectively avoiding instruction omission or incorrect execution, improving the accuracy and timeliness of trading responses, and significantly enhancing the operation efficiency of traders and the market response ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0018] Figure 1 This is a schematic diagram of the modules of the personal voice intelligent terminal system for the fund trading room of the present invention. Detailed implementation manners

[0019] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the example embodiments to those skilled in the art.

[0020] The present invention provides a personal voice intelligent terminal system for a fund trading room as Figure 1 shown, including a voice acquisition module, a voice preprocessing and structuring module, a voice feature analysis and urgency evaluation module, an intelligent recognition and evaluation module, and an adaptive semantic segmentation and priority regulation module: The voice acquisition module, through a real-time voice acquisition system, quickly captures each voice command of the trader; The voice acquisition system is responsible for extracting information from the trader's voice and quickly converting it into data for further processing. In high-frequency trading environments such as fund trading rooms, traders may need to quickly issue multiple commands in a short period of time, involving market trading, data query, or other operations. To ensure that no command is missed or delayed in processing, the system must be highly responsive and have low latency, capable of seamlessly processing continuous voice inputs and promptly recording each command. In addition, the system also needs to have strong anti-interference capabilities in noisy environments, able to accurately identify voices and eliminate background noise, ensuring that all voice commands can be quickly and accurately captured and transmitted to the backend processing system. The implementation of such a system is crucial for ensuring the efficiency and accuracy of the financial trading process, especially in scenarios where traders need to make quick decisions.

[0021] To obtain traders' voice commands in real time, various real-time voice acquisition systems can be adopted to ensure efficient and accurate capture of commands even in noisy environments. Common real-time voice acquisition systems include: microphone array systems, which work with multiple microphones to locate and clearly capture voice signals; super-directional microphones, which can focus on the trader's voice and suppress background noise to ensure accurate acquisition in complex environments; wireless microphone systems, which are suitable for traders in flexible and dynamic environments and transmit voice data wirelessly; integrated voice recognition headsets, which combine noise reduction technology and are suitable for traders to wear for long periods for high-frequency voice command input; intelligent speakers, which integrate voice acquisition and processing functions and can respond quickly to voice commands; near-field voice sensors, which are specially designed to accurately capture voice signals at short distances; voice activation systems, which automatically start voice recognition through keywords or trigger commands; and high-precision voice recognition software, which can process traders' voices in real time and convert them into commands. These systems use different technical means, such as noise suppression, echo cancellation, and sensitivity adjustment, to ensure accurate real-time capture of every voice command in high-noise and rapidly changing trading environments, thereby improving the efficiency and accuracy of traders' operations.

[0022] The voice preprocessing and structuring module preprocesses the acquired voice signals and converts the preprocessed voice information into a structured data set; The trading room environment usually has a high noise level, and background sounds such as keyboard clicks and phone calls may interfere with the recognition of voice input. Therefore, voice signals need to be processed through noise filtering, echo suppression, and volume normalization to ensure the clarity and accuracy of voice input during text conversion. In addition, audio signal enhancement can be performed on the voice during the processing to ensure that the voice signal can still be effectively recognized by the system even in complex environments. This preprocessing step is crucial for subsequent accurate analysis and is a prerequisite for improving voice recognition performance.

[0023] Key information extracted from traders' voice input, such as voice content, voice duration, pause length, voice pitch, speech rate, etc., can be organized into a data set. This data can be stored in a structured form for further analysis. The structured data set provides a clear basis for subsequent feature extraction, model training, and analysis. Through this data set, the system can identify which voice commands belong to emergency instructions, which instructions have a high level of complexity, and the overall trend of voice input, providing support for intelligent evaluation.

[0024] The specific steps for preprocessing the collected voice signals and converting them into a structured data set are as follows: First, noise suppression and echo cancellation technologies are used to remove background noise and reflected sounds to ensure the clarity of the voice signals. Next, the voice signals are enhanced and normalized to increase their volume and quality, making them more accurate in subsequent analysis. Subsequently, audio features of the voice, such as spectrum, speech rate, pitch, etc., are extracted through feature extraction algorithms (such as Mel Frequency Cepstral Coefficients MFCC), and the voice signals are segmented into multiple time frames, further divided into logical paragraphs, and pauses and tone changes are analyzed. Finally, the system converts the extracted voice features and information into a structured data set and performs annotation (such as instruction type, voice complexity, etc.) to provide an accurate data basis for subsequent voice recognition, instruction parsing, and intelligent decision-making.

[0025] The voice feature analysis and urgency assessment module extracts key features from the data set that reflect the rapidity and urgency of the trader's input instructions, comprehensively analyzes the extracted key features, and quantifies the voice urgency level. Extract key features from the data set that reflect the rapidity and urgency of the trader's input instructions. The extracted features include the comprehensive change in the pause frequency and pause duration in the voice and the fluctuation amplitude of the voice volume. The comprehensive change in the pause frequency and pause duration in the voice and the fluctuation amplitude of the voice volume are comprehensively analyzed under the detection window to generate a pause density reference value and a voice volume fluctuation reference value respectively, and the voice urgency level is quantified through the pause density reference value and the voice volume fluctuation reference value.

[0026] Frequent and short pauses in the voice can usually be regarded as one of the important manifestations of the rapidity and urgency of the trader's input instructions. The fundamental reason is that when traders are in a high-intensity and high-pressure trading environment, such as during sharp market fluctuations or near critical operation windows, they often need to issue multiple operation instructions within an extremely short time to ensure the timeliness of trading responses and the capture of market opportunities. In such a situation, the speech rhythm of traders significantly accelerates, and the transition time between each instruction is greatly compressed, forming a "short and frequent" pause pattern. This speech behavior shows the density and continuity of the instruction rhythm, reflecting the rapid thinking and operation needs of traders in a highly urgent state. In contrast, the voice instructions in the normal state usually have a gentle rhythm and natural pauses between instructions. Therefore, the voice feature of "frequent but extremely short pauses" can be used as one of the key signs of the urgency of voice input, indicating that the trader is quickly issuing multiple operation decisions, and the system needs to give priority to recognizing and responding to such high-priority voice inputs.

[0027] The specific steps for comprehensively analyzing the comprehensive change in the pause frequency and pause duration in the voice under the detection window to generate a pause density reference value are as follows: Within the detection window, identify and record all pause events in the speech. Each pause event consists of the speech duration before the pause and the pause duration, forming a sequence of pause events. Calculate the transition intensity factor for each pause event to measure the difference between the speech pause interval and the change in speech duration. The calculation expression for the transition intensity factor is as follows: , where is the transition intensity factor, used to measure the tension or change intensity between the th pause in the speech and the speech durations before and after it, is the th speech duration before the pause, that is, indicating the "talking time" or speech emission time between two pauses, is the th pause duration, representing the time interval from the stop of the speech to the start of the next speech segment, is a small constant, usually set to , used to avoid division-by-zero errors, is the tension amplification factor, used to adjust the sensitivity of the difference in speech duration between pauses, is the hyperbolic tangent function, and its output range is ; This step can accurately reflect the urgency and rapidity characteristics in the trader's voice commands by calculating the transition intensity factor for each pause, especially for the frequent and urgent voice inputs during the trading decision-making process.

[0028] Based on all the obtained transition intensity factors , generate a pause density reference value to quantify the urgency of the voice command. The formula for generating the pause density reference value is as follows: , where is the pause density reference value, is the natural base, is the total number of effective pauses.

[0029] This step can accurately evaluate the rapid commands issued by the trader in emergency situations by calculating the pause density reference value , especially applicable to trading scenarios that require quick decision-making and real-time response.

[0030] The larger the reference value of pause density generated by comprehensively analyzing the comprehensive changes in the pause frequency and pause duration in speech under the detection window, the more pauses there are in the trader's speech per unit time and the shorter the duration of each pause. This indicates that the trader is issuing multiple instructions quickly and continuously, and the transition between each instruction is very compact, showing an obvious emergency feature. Therefore, the higher the reference value of pause density, the faster the rhythm of the trader's input instructions and the more urgent the operation intention, and the instructions have a high degree of urgency; conversely, if the reference value of pause density is small, it means that the trader has a slow speaking speed and a long pause time when issuing instructions, and the instruction rhythm is loose, usually indicating that the current instruction does not have a high degree of urgency.

[0031] A drastic fluctuation in speech volume usually indicates that the trader is inputting instructions quickly and urgently. The reason is that in a high-pressure or urgent trading environment, traders often show emotional fluctuations, and this emotional reaction directly affects their speech output. Drastic fluctuations in speech volume are common in emotional states such as anxiety, tension, or excitement, and these emotional reactions often occur in situations of rapid decision-making or urgent tasks. When traders face market fluctuations and need to react quickly, their speech volume may increase, especially when issuing urgent trading instructions, and the change in speech volume is often spontaneous and non-linear. For example, a trader may quickly increase the volume at the beginning of an instruction to emphasize an urgent operation, and then quickly decrease the volume in the second half of the instruction, reflecting their eager and rapid decision-making process. The drastic fluctuation in volume is not only an external manifestation of emotion and urgency but also reflects their attempt to convey urgent information through a louder voice to ensure that the instruction can be executed quickly and accurately. Therefore, the degree of drastic fluctuation and frequent change in volume are an important sign of the urgency of an instruction and can effectively indicate the trader's mental state and reaction speed when facing urgent trading.

[0032] The specific steps for comprehensively analyzing the fluctuation range of speech volume under the detection window to generate a reference value of speech volume fluctuation are as follows: Within the detection window, obtain the energy value of the trader's speech input signal and establish an input signal energy sequence , , where is the energy value of the th frame of the speech signal, is the total number of speech frames. For the energy difference between adjacent frames, define the local mutation amplitude. The calculation expression of the local mutation amplitude is as follows: , where is the local mutation amplitude, that is, the magnitude of the volume change between adjacent frames (the th frame and the th frame), is the The energy value of the frame speech signal, It is the fluctuation sensitivity amplification factor, which is used to control the response intensity of the system to the volume change amplitude; By calculating the energy variation between adjacent speech frames, we can capture the instantaneous fluctuations in speech and identify the mutation features in speech signals. The impact of the violent fluctuations is enhanced, highlighting the intensity of voice changes when traders issue instructions in emergency situations.

[0033] Based on the local mutation amplitude obtained , further construct the speech volume fluctuation reference value, which is used to comprehensively measure the volume fluctuation intensity and rhythm density of the speech signal in the detection window. The calculation formula of the speech volume fluctuation reference value is as follows: , where is the reference value of voice volume fluctuation, is the rhythm reinforcement coefficient, which is used to adjust the response weight of continuous mutation behavior, usually , It is a continuous fluctuation frequency function, which is used to measure the number of continuous fluctuations in the current speech signal. Whether the frame is part of a continuous high fluctuation segment and identifies the position of the frame in the continuous fluctuation segment.

[0034] when The larger the value, the more violent the fluctuation of the trader's voice energy in the unit detection window, and the violent fluctuation shows obvious continuity and density. It can be judged that the instructions carried by the current voice input have higher urgency and time pressure. This reference value combines the local mutation characteristics and overall rhythm trend of the voice signal, and can effectively describe the voice change behavior of traders when issuing rapid instructions in emergency scenarios.

[0035] The larger the voice volume fluctuation reference value generated after comprehensive analysis of the fluctuation amplitude of the voice volume under the detection window, the more drastic the change of the trader's voice volume, reflecting that the trader's emotions fluctuate significantly and his tone is hasty when issuing instructions, which is usually accompanied by the rapid issuance of instructions and the urgency of decision-making. Therefore, the larger the voice volume fluctuation reference value, the more it can indicate that the instructions currently input by the trader are fast and urgent. On the contrary, if the monitored voice volume fluctuation reference value performance value is small, it means that the trader's voice remains relatively stable and the fluctuation amplitude is small, indicating that the trader is relatively calm and slow when inputting instructions, and the instructions are less urgent.

[0036] The intelligent identification and evaluation module inputs the analyzed key features into a pre-trained machine learning model and uses the model to intelligently evaluate the trader's instructions to determine whether the instructions are fast and urgent; Input the analyzed pause density reference value and voice volume fluctuation reference value into a pre-trained machine learning model. Generate an instruction urgency coefficient through the machine learning model, and use the instruction urgency coefficient to intelligently evaluate the trader's instruction, so as to determine whether the instruction has the characteristics of being fast and urgent.

[0037] The pre-trained machine learning model mentioned here refers to a mathematical model with recognition and discrimination capabilities formed through systematic training based on a large number of historical voice input samples in the application scenario of trader voice data. By learning the relationship between the change in pause density, the feature of volume fluctuation in the historical voice instructions of traders and the actual instruction urgency level, the model automatically establishes a mapping relationship between the features and the output. During the training process, the system will input thousands of labeled training data, which not only include input features such as pause density reference values and voice volume fluctuation reference values, but also include the urgency labels corresponding to each sample (such as "high urgency", "medium urgency", "not urgent", etc.). The model gradually learns how to accurately predict the instruction urgency according to the combination of input features by repeatedly iterating and optimizing internal parameters (such as weight matrices, activation function coefficients, etc.). The essence of this "pre-training" process is to enable the model to extract potential patterns and rules from a large amount of existing data before being formally put into real-time application, forming an inference and judgment ability that can widely adapt to different traders, different environments, and different speech rate characteristics.

[0038] Once the pre-training is completed, the machine learning model can directly receive new input data (i.e., the pause density reference value and voice volume fluctuation reference value collected in real time) during the real-time application stage, and quickly output an instruction urgency coefficient. This coefficient can be a continuous value (such as a decimal between 0 and 1, the higher the value, the more urgent), or a classification label (such as a binary classification of "urgent / non-urgent"). The model infers whether the trader's current instruction has the characteristics of being fast and urgent through this kind of reasoning, realizing intelligent and automated real-time evaluation. Compared with the traditional static rule judgment method, using a pre-trained machine learning model has higher flexibility and generalization ability, and can adapt to the differences in the natural language behaviors of traders under different market environments, emotional fluctuations or speech rate changes, thus significantly improving the accuracy and response speed of instruction urgency recognition. This mechanism ensures that in scenarios such as high-frequency trading and sudden market changes that require extremely fast decision-making, the system can assist trading operations in a highly reliable manner.

[0039] The machine learning model is not limited here, and it can realize the comprehensive analysis of the pause density reference value and the voice volume fluctuation reference value to generate an instruction urgency coefficient Any machine learning model can be used. To implement the technical solution of the present invention, the present invention provides a specific implementation method: Instruction urgency coefficient The generation formula is as follows; , where, and are respectively the preset proportionality coefficients of the pause density reference value and the voice volume fluctuation reference value , and and are both greater than 0.

[0040] The "preset proportionality coefficient" refers to two weight parameters and introduced in the formula, which respectively correspond to the pause density reference value and the voice volume fluctuation reference value in calculating the instruction urgency coefficient . The so-called "preset" means that these two parameters are not calculated in real time through the formula, but are preset according to experience or statistical laws during the model design stage. The purpose of setting these two proportionality coefficients is to reflect the influence intensity and contribution degree of different physical quantity dimension features (such as pause frequency and volume fluctuation) on the final evaluation result when integrating them.

[0041] Specifically, is used to control the influence weight of the pause density reference value on the instruction urgency coefficient , while controls the influence weight of the voice volume fluctuation reference value . If in a certain application scenario, the pause density is more sensitive to the discrimination of instruction urgency, then can be set, and vice versa. Such a design makes the model have a certain degree of adjustability and scene adaptability. At the same time, setting these two parameters can also play a role in numerical balance, preventing one item from having too large a dominant effect on the input of the logarithmic function and affecting the stability of the urgency coefficient. Therefore, the role of the preset proportionality coefficient not only lies in expressing feature contributions, but also has mathematical meanings such as smoothing, multi-dimensional normalization, and model adjustment flexibility.

[0042] From the instruction urgency coefficient, it can be seen that the larger the reference value of the pause density generated by comprehensively analyzing the comprehensive changes in the pause frequency and pause duration in the voice under the detection window, and the larger the reference value of the voice volume fluctuation generated by comprehensively analyzing the fluctuation amplitude of the voice volume under the detection window, the larger the instruction urgency coefficient generated when the trader's instructions are intelligently evaluated by a pre-trained machine learning model, indicating that the probability that the instructions issued by the trader are fast and urgent is greater. Conversely, it indicates that the probability that the instructions issued by the trader are fast and urgent is smaller.

[0043] Compare and analyze the instruction urgency coefficient generated when the trader's instructions are intelligently evaluated by a pre-trained machine learning model with the pre-set reference threshold of the instruction urgency coefficient to determine whether the instruction has the characteristics of being fast and urgent. The judgment logic is as follows: If the instruction urgency coefficient is greater than the pre-set reference threshold of the instruction urgency coefficient, it is determined that the instruction has the characteristics of being fast and urgent, indicating that the instruction issued by the trader is fast and urgent; If the instruction urgency coefficient is less than or equal to the pre-set reference threshold of the instruction urgency coefficient, it is determined that the instruction does not have the characteristics of being fast and urgent, indicating that the instruction issued by the trader does not have urgency.

[0044] The adaptive semantic segmentation and priority regulation module, when it recognizes that the instructions issued by the trader are fast and urgent, divides the continuous voice input of the trader into multiple paragraphs, and dynamically determines the parsing priority of each paragraph of voice in combination with the pauses, accents and context information of the current voice. When it recognizes that the instructions issued by the trader are fast and urgent, the voice instruction is parsed by combining the deep neural network (DNN) and the adaptive segmentation technology. Its core function is to improve the understanding ability and response accuracy for continuous and high-frequency voice input, ensuring that the system can achieve semantic integrity, timely response and accurate execution recognition processing in a multi-instruction and high-pressure environment. In actual application scenarios, traders often quickly issue multiple trading-related instructions in a very short time, such as "Immediately buy 1,000 shares of Company A, and then cancel the previous order". Such instructions often have a fast speech rate and lack obvious pauses between instructions, and are easily misjudged or merged by traditional voice recognition systems, resulting in execution errors.

[0045] By introducing the DNN model, the system can perceive subtle changes in speech (such as sudden changes in speech rate, heightened emotion, and rising intonation) based on its deep feature learning ability, thereby accurately judging the logical structure of the speech content. The adaptive segmentation technology can reasonably divide a continuous speech into multiple speech segments with independent semantics according to changes such as pauses, accents, and semantic turns in the actual speech data. The system further combines the context to dynamically judge the parsing priority of each speech segment, enabling important and urgent instructions to be recognized and processed first, while non-critical content is parsed later, thus achieving an intelligent scheduling of "parsing by priority on demand".

[0046] When it is recognized that the instructions issued by the trader are fast and urgent, the specific steps for dividing the continuous speech input of the trader into multiple segments and dynamically determining the parsing priority of each speech segment in combination with the pauses, accents, and context information of the current speech are as follows: When it is recognized that the instructions issued by the trader are fast and urgent, by analyzing the pauses, accents, and context information in the speech, the speech input of the trader is dynamically divided into multiple speech segments, and the parsing priority of each segment is adjusted according to the speech features. The pause time, accent intensity, and context relevance in the speech will directly affect the urgency of the instructions. The priority of each segment is dynamically determined based on these factors, and the priority adjustment formula is as follows: , where is the priority score assigned to each speech segment, and this score will determine the priority of the speech instruction in parsing and execution. In the speech input, is the pause duration of the th speech segment, is the intensity of the accented part of the th speech segment, is the degree of association between the context information of the th speech segment and the current trading environment, is the weight coefficient of the pause time, which adjusts the influence degree of the pause time on the priority adjustment, is the weight coefficient of the accent intensity, which adjusts the influence degree of the accent intensity on the priority adjustment, is the weight coefficient of the context relevance, which adjusts the influence degree of the context relevance on the priority adjustment;

[0047] After adjusting the priority of each speech segment based on the speech features, the priorities of all segments are sorted to ensure that the instructions are executed according to the preset priority, so as to ensure that the urgent instructions can be processed first. The priority sorting formula is as follows: , where is the sorting position of the th voice segment during the final execution, is the priority of the th voice segment, that is, the priority score assigned to the th voice segment, is the total number of voice segments.

[0048] By prioritizing each voice segment, the system can ensure that high-priority instructions are executed first, reducing latency and errors. This mechanism ensures that the trader's emergency instructions are responded to in a timely manner, avoiding potential losses caused by instruction delays.

[0049] The present invention captures and preprocesses voice signals in real time, extracts and quantifies the urgency of the trader's voice instructions, and determines the priority of the instructions through an intelligent evaluation model, so as to ensure that emergency instructions can be preferentially recognized and processed. Further, through an adaptive segmentation technique and context analysis, the system can dynamically adjust the priority of voice parsing according to information such as pauses and accents in the trader's voice, effectively avoiding instruction omission or incorrect execution, improving the accuracy and timeliness of trading responses, and significantly enhancing the trader's operation efficiency and market response ability.

[0050] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0051] Only some exemplary embodiments of the present invention have been described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0052] It should be noted that in this article, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0053] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0054] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0055] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0056] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0057] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0058] As described above, this is only the specific implementation of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.

[0059] Only some exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.

Claims

1. A personal voice intelligent terminal system for a fund trading room, characterized in that, It includes a voice acquisition module, a voice pre - processing and structuring module, a voice feature analysis and urgency assessment module, an intelligent recognition and assessment module, and an adaptive semantic segmentation and priority regulation module: The voice acquisition module, through a real - time voice acquisition system, quickly captures each voice command of the trader; The voice pre - processing and structuring module pre - processes the collected voice signals and converts the pre - processed voice information into a structured data set; The voice feature analysis and urgency assessment module extracts key features from the data set that reflect the rapidity and urgency of the trader's input commands, comprehensively analyzes the extracted key features, and quantifies the voice urgency level; The intelligent recognition and assessment module inputs the analyzed key features into a pre - trained machine learning model, uses the model to intelligently assess the trader's commands, and thus determines whether the commands have the characteristics of being rapid and urgent; The adaptive semantic segmentation and priority regulation module, when it recognizes that the commands issued by the trader are rapid and urgent, divides the continuous voice input of the trader into multiple paragraphs, and dynamically determines the parsing priority of each paragraph of voice in combination with the current voice pauses, accents, and context information; Extract key features from the data set that reflect the rapidity and urgency of the trader's input commands. The extracted features include the comprehensive change in the pause frequency and pause duration in the voice and the fluctuation amplitude of the voice volume. The comprehensive change in the pause frequency and pause duration in the voice and the fluctuation amplitude of the voice volume are comprehensively analyzed under the detection window to generate a pause density reference value and a voice volume fluctuation reference value respectively, and the voice urgency level is quantified through the pause density reference value and the voice volume fluctuation reference value.

2. The personal voice intelligent terminal system for a fund trading room according to claim 1, wherein The specific steps for pre - processing the collected voice signals and converting them into a structured data set are as follows: Adopt noise suppression and echo cancellation technologies to remove background noise and reflected sounds to ensure the clarity of the voice signals; The voice signals are enhanced and standardized to improve their volume and quality, making them more accurate in subsequent analysis; Extract the audio features of the voice through a feature extraction algorithm, segment the voice signals into multiple time frames, further divide them into logical paragraphs, and analyze the pauses and tone changes; Convert the extracted voice features and information into a structured data set and perform annotation to provide an accurate data basis for subsequent voice recognition, command parsing, and intelligent decision - making.

3. The personal voice intelligent terminal system for the fund trading room according to claim 1, characterized in that, The specific steps for comprehensively analyzing the comprehensive change in the pause frequency and pause duration in the voice under the detection window to generate a pause density reference value are as follows: Within the detection window, identify and record all pause events in the voice. Each pause event consists of the voice duration before the pause and the pause duration, forming a pause event sequence. Calculate the transition intensity factor of each pause event to measure the difference between the voice pause interval and the change in voice duration. The calculation expression of the transition intensity factor is as follows: , where is the transition intensity factor, is the speech duration before the -th pause, is the duration of the -th pause, is a small constant, is the tenseness amplification factor, is the hyperbolic tangent function; Based on all the obtained transition intensity factors , a pause density reference value is generated to quantify the urgency of the voice command. The formula for generating the pause density reference value is as follows: , where is the reference value of pause density, is the natural base, is the total number of effective pauses.

4. The personal voice intelligent terminal system for the fund trading room according to claim 1, characterized in that, The specific steps for comprehensively analyzing the fluctuation amplitude of the voice volume under the detection window to generate a voice volume fluctuation reference value are as follows: Within the detection window, obtain the energy value of the trader's voice input signal and establish an input signal energy sequence , , where is the energy value of the th frame of the voice signal, is the total number of voice frames. For the energy difference between adjacent frames, define the local mutation amplitude. The calculation expression of the local mutation amplitude is as follows: , where is the local mutation amplitude, is the energy value of the nth frame of voice signal, is the fluctuation sensitivity amplification factor; Based on the obtained local mutation amplitude , a reference value of voice volume fluctuation is further constructed to comprehensively measure the volume fluctuation intensity and rhythm density of the voice signal within the detection window. The calculation formula of the reference value of voice volume fluctuation is as follows: , where is the reference value of voice volume fluctuation, is the rhythm enhancement coefficient, is the continuous fluctuation times function.

5. The personal voice intelligent terminal system for a fund trading room according to claim 1, wherein The analyzed pause density reference value and voice volume fluctuation reference value are input into a pre-trained machine learning model. The machine learning model generates an instruction urgency coefficient, and the trader's instruction is intelligently evaluated through the instruction urgency coefficient, so as to determine whether the instruction has the characteristics of being fast and urgent.

6. The personal voice intelligent terminal system for a fund trading room according to claim 5, characterized in that The instruction urgency coefficient generated when the trader's instruction is intelligently evaluated by the pre-trained machine learning model is compared and analyzed with the pre-set reference threshold of the instruction urgency coefficient to determine whether the instruction has the characteristics of being fast and urgent. The judgment logic is as follows: If the instruction urgency coefficient is greater than the pre-set reference threshold of the instruction urgency coefficient, it is determined that the instruction has the characteristics of being fast and urgent, indicating that the instruction issued by the trader is fast and urgent; If the instruction urgency coefficient is less than or equal to the pre-set reference threshold of the instruction urgency coefficient, it is determined that the instruction does not have the characteristics of being fast and urgent, indicating that the instruction issued by the trader does not have urgency.

7. The personal voice intelligent terminal system for a fund trading room according to claim 6, wherein When it is recognized that the instruction issued by the trader is fast and urgent, the continuous voice input of the trader is divided into multiple paragraphs, and the specific steps for dynamically determining the parsing priority of each paragraph of voice in combination with the pauses, accents and context information of the current voice are as follows: When it is recognized that the instruction issued by the trader is fast and urgent, the voice input of the trader is dynamically divided into multiple voice paragraphs by analyzing the pauses, accents and context information in the voice, and the parsing priority of each paragraph is adjusted according to the voice characteristics. The priority adjustment formula is as follows: , where is the priority score assigned to each speech segment, is the pause duration of the th speech segment in the speech input, is the intensity of the stressed part of the th speech segment, is the degree of relevance between the context information of the th speech segment and the current transaction environment, is the weight coefficient of the pause time, is the weight coefficient of the stress intensity, is the weight coefficient of the context relevance; After adjusting the priority of each voice paragraph based on the voice characteristics, the priorities of all paragraphs are sorted to ensure that the instructions are executed according to the preset priority, so as to ensure that the urgent instructions can be processed first. The priority sorting formula is as follows: , where is the sorting position of the th voice segment during final execution, is the priority of the th voice segment, is the total number of voice segments.

Citation Information

Patent Citations

  • Personal voice intelligent terminal for fund transaction room

    CN119559938A

  • Telephone answering system based on AI

    CN119854414A

  • Method for recogning emergency speech using gmm

    KR1020120130371A

Cited By

  • AI-based mediation record law fact keyword extraction system and method

    CN121029988A