Multilingual real-time translation cash register terminal system

By using the interactive language recognition, ambiguous content recognition, and alternative translation filtering modules of the multilingual real-time translation POS terminal system, the problem of inaccurate translation during the checkout process has been solved, enabling more efficient and accurate multilingual interaction and improving the transaction experience.

CN120046627BActive Publication Date: 2026-05-01SHENZHEN ACME TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN ACME TECH CO LTD
Filing Date
2025-01-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing translation systems struggle to accurately handle complex multilingual interactions during checkout, especially when dealing with technical terms and ambiguous words, leading to misunderstandings and inaccurate translations, which impacts transaction efficiency and customer experience.

Method used

The multilingual real-time translation POS terminal system includes an interactive language recognition module, a translation and ambiguous content recognition module, and a candidate translation filtering module. It quickly identifies the languages ​​of customers and cashiers, identifies ambiguous words and sentences, and uses a multi-semantic weight information database based on the adjacent context and POS scenario to filter out the final translation result.

Benefits of technology

It improves translation accuracy and communication efficiency in multilingual interactions, provides a more convenient and smooth service experience, reduces ambiguity, and enhances the accuracy and efficiency of the checkout process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046627B_ABST
    Figure CN120046627B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of language translation, and particularly discloses a multilingual real-time translation cash register terminal system, which comprises an interactive language recognition module, a translation and ambiguity content recognition module, and an alternative translation screening module. The interactive language recognition module judges the interactive language of a customer based on the first interactive voice data input by the customer. The translation and ambiguity content recognition module translates the interactive voice data input by the customer or a cashier in real time and determines all ambiguous words, all ambiguous sentences and all alternative translation results in the corresponding interactive voice data based on the interactive language of the customer or the interactive language of the cashier. The alternative translation screening module screens the final translation result of the real-time input interactive voice data from all the alternative translation results based on the adjacent context of all the ambiguous words, the adjacent context of all the ambiguous sentences and a preset cash register scenario multi-word meaning weight information base. The application improves the communication efficiency and accuracy in the cash register process and provides more convenient and smooth service experience for the customer and the cashier.
Need to check novelty before this filing date? Find Prior Art

Description

Multilingual real-time translation POS terminal system Technical Field

[0001] This invention relates to the field of language translation technology, and in particular to a multilingual real-time translation POS terminal system. Background Technology

[0002] In a globalized economy, business activities are becoming increasingly frequent and diversified, with ever-increasing communication and transactions between people from different countries and regions. In the retail sector, the checkout process is a crucial step in completing a transaction, but language barriers can cause inconvenience for both customers and cashiers, impacting transaction efficiency and customer experience. Therefore, the need for multilingual real-time translation checkout terminal systems has emerged. With the development of speech recognition and machine translation technologies, multilingual communication has become possible.

[0003] While some businesses have begun using simple translation tools or devices, these often have limited functionality and struggle to meet the demands of complex checkout scenarios. The checkout process involves various technical terms and specific contexts related to product information, prices, and payment methods, requiring more accurate and intelligent translation systems. Existing translation systems may perform well in everyday conversations, but in the specific context of checkout, misunderstandings or inaccurate translations may occur. Furthermore, different languages ​​differ significantly in grammar, vocabulary, and expression; existing translation systems often fail to accurately grasp the true meaning of ambiguous words or sentences.

[0004] Therefore, this invention proposes a multilingual real-time translation POS terminal system. Summary of the Invention

[0005] This invention provides a multilingual real-time translation POS terminal system. Utilizing an interactive language recognition module, it can quickly and accurately determine the customer's interactive language, laying the foundation for subsequent translation work and improving the relevance and efficiency of the service. The translation and ambiguous content recognition module can translate interactive voice data and identify ambiguous words, sentences, and alternative translation results, comprehensively considering the complexities of translation. The alternative translation filtering module filters based on the adjacent context of ambiguous words and sentences and a multi-semantic weight information database of the POS scenario, ensuring that the final translation result is more consistent with the actual POS scenario, improving the accuracy and practicality of the translation. This system effectively solves translation problems in multilingual interactions, reduces ambiguity, improves communication efficiency and accuracy during the POS process, and provides a more convenient and seamless service experience for customers and cashiers.

[0006] This invention provides a multilingual real-time translation POS terminal system, comprising:

[0007] The interactive language recognition module is used to determine the customer's interactive language based on the initial interactive voice data input by the customer;

[0008] The translation and ambiguous content recognition module is used to translate the interactive voice data input in real time by customers or cashiers based on the customer's or cashier's interactive language, and to identify all ambiguous words, all ambiguous sentences, and all alternative translation results in the corresponding interactive voice data.

[0009] The alternative translation filtering module is used to filter the final translation result of the real-time input interactive voice data from all the alternative translation results based on the neighboring contexts of all ambiguous words and all ambiguous sentences in the real-time input interactive voice data, as well as the preset multi-semantic weight information database of the cashier scenario.

[0010] Preferably, the interactive language recognition module includes:

[0011] The spectrum generation submodule is used to divide the speech signal of the customer's first interactive speech data based on the preset frame length, frame shift and moving window function, to obtain the windowed speech signal of each frame of the first interactive speech data, and to perform a fast Fourier transform on the windowed speech signal of each frame of the first interactive speech data to obtain the frequency domain spectrum of the speech signal of each frame of the first interactive speech data.

[0012] The resonance feature periodicity analysis submodule is used to analyze the frequency domain spectrum of each frame of the speech signal in the first interactive speech data to determine the resonance period and periodicity representative resonance change characteristics of the first interactive speech data.

[0013] The high-energy temporal distribution feature analysis submodule is used to analyze the high-energy temporal distribution features of the first interactive speech data based on the frequency domain spectrum of each frame of the speech signal.

[0014] The interactive language identification submodule is used to input the resonance period, periodic representative resonance change features, and high-energy time-domain distribution features of the initial interactive voice data into the language identification model to obtain the customer's interactive language.

[0015] Preferably, the resonance characteristic periodicity analysis submodule includes:

[0016] The resonance feature change trend curve generation unit is used to count the number of resonance peaks and the frequency points of resonance peaks in the spectrum of each frame of the first interactive speech data, and to perform time-series fitting on the number of resonance peaks and the frequency points of resonance peaks in the spectrum of all frames of the first interactive speech data to obtain the resonance peak change trend curve and the resonance peak frequency point change trend curve of the first interactive speech data.

[0017] The resonance period analysis unit is used to analyze the variation period of the formant variation trend curve and the variation period of the formant frequency point variation trend curve of the first interactive voice data based on the autocorrelation analysis method, and takes the least common multiple of the variation period of the formant variation trend curve and the variation period of the formant frequency point variation trend curve of the first interactive voice data as the resonance period of the first interactive voice data.

[0018] The periodic representative resonance variation feature analysis unit is used to input the formant variation trend curve, the formant frequency point variation trend curve, and the resonance period of the first interactive speech data into the periodic feature generalization model to obtain the periodic representative resonance variation features of the first interactive speech data.

[0019] The preferred high-energy time-domain distribution characteristic analysis submodule includes:

[0020] The frequency spectrum centroid change trend curve fitting unit is used to determine the frequency spectrum centroid of each frame of the speech signal based on the frequency domain spectrum of each frame of the first interactive speech data, and to perform time-series fitting on the frequency spectrum centroid of all frames of the first interactive speech data to obtain the frequency spectrum centroid change trend curve of the first interactive speech data.

[0021] The energy concentration frequency band change trend curve fitting unit is used to simultaneously divide the frequency band and analyze the frequency band energy distribution characteristics of the frequency domain spectrum of each frame of the first interactive voice data to obtain the energy concentration frequency band change trend curve of the first interactive voice data.

[0022] The merged feature analysis unit is used to input the spectral centroid change trend curve and the energy concentration frequency band change trend curve of the first interactive speech data into the merged feature analysis model to obtain the high-energy time-domain distribution characteristics of the first interactive speech data.

[0023] Preferably, the energy concentration frequency band variation trend curve fitting unit includes:

[0024] The frequency band division and frequency band energy calculation subunit is used to simultaneously divide the frequency domain spectrum of each frame of the first interactive voice data into frequency bands, obtain multiple frequency bands of each frame of the first interactive voice data, and calculate the frequency band energy of each frequency band of each frame of the voice data.

[0025] The normalization and curve fitting subunit is used to normalize the frequency band energy of all the same frequency band in all frames of the first interactive voice data and then fit it according to the time sequence to obtain the energy change trend curve of each frequency band of the first interactive voice data.

[0026] The curve alignment analysis subunit is used to perform alignment analysis on the energy change trend curves of all frequency bands of the first interactive voice data, and to fit the energy concentration frequency band change trend curve of the first interactive voice data.

[0027] Preferably, the translation and ambiguous content recognition module includes:

[0028] The real-time customer voice translation submodule is used to convert the real-time interactive voice data into text by treating the customer's interactive language as the input language when it receives interactive voice data in real time, so as to obtain the customer's voice input text.

[0029] The customer ambiguous word retrieval submodule is used to take the cashier's interactive language as the output language and retrieve all the definitions of each word in the customer's voice input text in the output language from the multilingual lexicon. It also takes words in the customer's voice input text that contain more than one definition in the output language as ambiguous words in the corresponding interactive voice data.

[0030] The Customer Ambiguous Sentence Retrieval Submodule is used to perform grammatical deconstruction on each sentence in the customer's voice input text to obtain all the grammatical structures of each sentence, and to treat sentences in the customer's voice input text that contain more than one grammatical structure as ambiguous sentences in the corresponding interactive voice data.

[0031] The Customer Alternate Translation Result Generation Submodule is used to generate all alternative translation results for the corresponding interactive voice data in the corresponding output language, based on all ambiguous words and sentences in the interactive voice data input by the customer in real time, as well as the customer's voice input text.

[0032] Preferably, the translation and ambiguous content recognition module includes:

[0033] The cashier's real-time voice translation submodule is used to convert the real-time interactive voice data into text when the cashier's interactive language is received in real time, so as to obtain the cashier's voice input text.

[0034] The cashier ambiguous word retrieval submodule is used to take the customer's interaction language as the output language and retrieve all the definitions of each word in the cashier's voice input text in the output language from the multilingual lexicon. Words in the cashier's voice input text that contain more than one definition in the output language are treated as ambiguous words in the corresponding interactive voice data.

[0035] The Cashier Ambiguity Retrieval Submodule is used to perform grammatical deconstruction on each sentence in the cashier's voice input text to obtain all the grammatical structures of each sentence, and to treat sentences in the cashier's voice input text that contain more than one grammatical structure as ambiguous sentences in the corresponding interactive voice data.

[0036] The cashier alternative translation result generation submodule is used to generate all alternative translation results for the corresponding interactive voice data in the corresponding output language based on all ambiguous words and sentences in the interactive voice data input by the cashier in real time, as well as the cashier's voice input text.

[0037] Preferred alternative translation filtering modules include:

[0038] The word meaning filtering submodule is used to determine the actual meaning of each ambiguous word based on the nearby context of each ambiguous word in the real-time input interactive voice data and the preset multi-word meaning weight information database of the cashier scenario.

[0039] The grammatical structure filtering submodule is used to determine the weight value of each grammatical structure of each ambiguous sentence based on the neighboring context of each ambiguous sentence, and to take the grammatical structure with the largest weight value among all grammatical structures of each ambiguous sentence as the actual grammatical structure of each ambiguous sentence.

[0040] The translation result filtering submodule is used to filter the final translation result of the real-time input interactive voice data from all candidate translation results based on the actual meaning of all ambiguous words and the actual grammatical structure of all ambiguous sentences in the real-time input interactive voice data.

[0041] The preferred semantic filtering submodule includes:

[0042] The total weighting unit is used to determine the total weighting value of each meaning of each ambiguous word based on the nearby context of each ambiguous word in the real-time input interactive voice data and the preset multi-meaning weight information database of the cashier scenario.

[0043] The definition filtering unit is used to take the definition with the highest total weight among all the definitions of each ambiguous word as the actual meaning of each ambiguous word.

[0044] Preferred, the total weighted unit includes:

[0045] The first weighting subunit is used to determine the first weighting value for each interpretation of each ambiguous word based on the neighborhood context of each ambiguous word in real-time input interactive voice data.

[0046] The second weighting subunit is used to retrieve the preset multi-semantic weight information database for the cashier scenario and determine the second weighting value for each definition of each ambiguous word.

[0047] The total weighting subunit is used to take the sum of the first and second weighting values ​​of each definition of each ambiguous word as the total weighting value of each definition.

[0048] The beneficial effects of this invention compared to existing technologies are as follows: The interactive language recognition module can quickly and accurately determine the customer's interactive language, laying the foundation for subsequent translation work and improving the targeting and efficiency of the service. The translation and ambiguous content recognition module can translate interactive voice data and identify ambiguous words, sentences, and alternative translation results, comprehensively considering the complexities in translation. The alternative translation filtering module filters based on the adjacent context of ambiguous words and sentences and a multi-semantic weight information database for the checkout scenario, ensuring that the final translation result is more consistent with the actual checkout scenario, improving the accuracy and practicality of the translation. This system can effectively solve translation problems in multilingual interactions, reduce ambiguity, improve communication efficiency and accuracy during the checkout process, and provide customers and cashiers with a more convenient and seamless service experience.

[0049] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.

[0050] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0051] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0052] Figure 1 is a schematic diagram of the functional modules contained in the multilingual real-time translation POS terminal system in an embodiment of the present invention.

[0053] Figure 2 is a schematic diagram of the functional sub-modules contained in the interactive language recognition module in an embodiment of the present invention;

[0054] Figure 3 is a schematic diagram of the functional sub-modules contained in the translation and ambiguous content recognition module in an embodiment of the present invention;

[0055] Figure 4 is a schematic diagram of the internal functional sub-modules of another translation and ambiguous content recognition module in an embodiment of the present invention;

[0056] Figure 5 is a schematic diagram of the functional sub-modules contained in the alternative translation screening module in an embodiment of the present invention. Detailed Implementation

[0057] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0058] Example 1:

[0059] This invention provides a multilingual real-time translation POS terminal system, as shown in Figure 1, including:

[0060] The interactive language recognition module is used to determine the customer's interactive language based on the initial interactive voice data input by the customer;

[0061] The translation and ambiguous content recognition module is used to translate the interactive voice data input in real time by customers or cashiers based on the customer's or cashier's interactive language, and to identify all ambiguous words, all ambiguous sentences, and all alternative translation results in the corresponding interactive voice data.

[0062] The alternative translation filtering module is used to filter the final translation result of the real-time input interactive voice data from all the alternative translation results based on the neighboring contexts of all ambiguous words and all ambiguous sentences in the real-time input interactive voice data, as well as the preset multi-semantic weight information database of the cashier scenario.

[0063] In this embodiment, the initial interactive voice data refers to the first set of voice data input by the customer when interacting with the POS terminal system. For example, a foreign customer says to the POS terminal for the first time, "Hello, I want to buy this item."

[0064] In this embodiment, the customer's interaction language refers to the language used by the customer when communicating with the POS terminal, such as English, French, Chinese, etc.

[0065] In this embodiment, interactive voice data refers to the voice information input by customers or cashiers during communication with a multilingual real-time translation POS terminal system.

[0066] In this embodiment, ambiguous words refer to words that have multiple different meanings in a specific linguistic context. For example, "bank" in English can mean both "bank" and "riverbank".

[0067] In this embodiment, an ambiguous sentence refers to a sentence that has multiple possible grammatical structures or semantic interpretations. For example, "He saw her smile" can be understood as "He saw her, and then she smiled," or it can be understood as "He saw the action of her smiling."

[0068] In this embodiment, all alternative translation results refer to the various possible translation results obtained after preliminary processing of a segment of interactive voice data.

[0069] In this embodiment, the context of an ambiguous word refers to the linguistic environment, such as the words and sentences surrounding the ambiguous word. For example, in the sentence "I went to the bank to withdraw money," the context surrounding "bank" helps to determine that it means "bank."

[0070] In this embodiment, the adjacent context of an ambiguous sentence refers to the surrounding content of the ambiguous sentence. For example, if the preceding sentence says "He saw her smile," and the following sentence says "He also became happy," the first interpretation is more likely.

[0071] In this embodiment, the preset multi-semantic weight information database for checkout scenarios refers to a pre-established set of weight information that assigns importance or frequency of different meanings to common words in checkout scenarios. For example, in a checkout scenario, the weight of the meaning of "charge" as "to charge" may be higher than that of "to charge".

[0072] In this embodiment, the final translation result refers to the most accurate translation result that best fits the current context after screening and confirmation.

[0073] The beneficial effects of the above technologies are as follows: The interactive language recognition module can quickly and accurately determine the customer's interactive language, laying the foundation for subsequent translation work and improving the targeting and efficiency of the service. The translation and ambiguous content recognition module can translate interactive voice data and identify ambiguous words, ambiguous sentences, and alternative translation results, comprehensively considering the complexities in translation. The alternative translation filtering module filters based on the adjacent context of ambiguous words and sentences and a multi-semantic weight information database of the checkout scenario, ensuring that the final translation result is more consistent with the actual checkout scenario, improving the accuracy and practicality of the translation. This system can effectively solve the translation problem in multilingual interaction, reduce ambiguity, improve communication efficiency and accuracy in the checkout process, and provide customers and cashiers with a more convenient and smooth service experience.

[0074] Example 2:

[0075] Based on Example 1, the interactive language recognition module, referring to Figure 2, includes:

[0076] The spectrum generation submodule is used to divide the speech signal of the customer's first interactive speech data based on the preset frame length, frame shift and moving window function, to obtain the windowed speech signal of each frame of the first interactive speech data, and to perform a fast Fourier transform on the windowed speech signal of each frame of the first interactive speech data to obtain the frequency domain spectrum of the speech signal of each frame of the first interactive speech data.

[0077] The resonance feature periodicity analysis submodule is used to analyze the frequency domain spectrum of each frame of the speech signal in the first interactive speech data to determine the resonance period and periodicity representative resonance change characteristics of the first interactive speech data.

[0078] The high-energy temporal distribution feature analysis submodule is used to analyze the high-energy temporal distribution features of the first interactive speech data based on the frequency domain spectrum of each frame of the speech signal.

[0079] The interactive language identification submodule is used to input the resonance period, periodic representative resonance change features, and high-energy time-domain distribution features of the initial interactive voice data into the language identification model to obtain the customer's interactive language.

[0080] In this embodiment, the preset frame length and frame shift refer to pre-set parameters used for processing the speech signal. The frame length is the duration of each frame when the speech signal is segmented, and the frame shift is the time interval between the start positions of two adjacent frames.

[0081] In this embodiment, the moving window function is a function used in framing speech signals to smooth frame boundaries and reduce spectral leakage. For example, a Hamming window function can be used.

[0082] In this embodiment, the voice signal of the customer's first interactive voice data is divided based on the preset frame length, frame shift, and moving window function to obtain the windowed voice signal of each frame of the first interactive voice data. This means that the voice signal of the customer's first input is divided into segments of frames according to the preset frame length and frame shift, and the moving window function is used to process each frame to obtain the windowed voice signal for subsequent analysis and processing.

[0083] In this embodiment, the resonance period and periodicity of the initial interactive voice data represent the resonance change characteristics: the resonance period refers to the time interval of the periodic regularity of the resonance of the voice signal in the initial interactive voice data; the periodicity represents the resonance change characteristics, which are the key features that can reflect this periodic resonance change characteristics.

[0084] In this embodiment, high-energy temporal distribution characteristics refer to the temporal distribution of the high-energy portions in the initial interactive speech data. For example, analysis may reveal that the speech signal has concentrated and strong energy during certain time periods, while its energy is weaker during other time periods. This distribution pattern and characteristic of energy on the time axis constitutes high-energy temporal distribution characteristics. It can help identify important parts of the speech, distinguish different speech patterns, and assist in language identification. For instance, in a piece of initial interactive speech data, the sound energy is very strong during the 0-0.5 second period, significantly weakens during the 0.5-1 second period, and then strengthens again during the 1-1.5 second period. This distribution of energy strength at different time points, such as the duration of concentrated energy frequency bands and the pattern of energy strength changes, constitutes high-energy temporal distribution characteristics. Another example: in another speech segment, the high-frequency energy is very high during the first 0-0.2 seconds, then relatively low and stable high-frequency energy and high mid-frequency energy during the next 0.2-0.8 seconds, and finally, a sudden increase in high-frequency energy during the last 0.8-1 seconds while mid-frequency energy returns to a stable level. This also reflects a high-energy temporal distribution characteristic.

[0085] In this embodiment, the language identification model is a trained model used to determine the language of speech based on its features. For example, with a preset frame length of 20 milliseconds and a frame shift of 10 milliseconds, a 1-second speech segment is divided using a Hanning window function. Analysis reveals that the resonance period of this speech segment is approximately 50 milliseconds. The periodicity represents the specific variation pattern of the formant frequency. The high-energy temporal distribution features are as follows: in the first 0-0.2 seconds, the high-frequency energy is very high; in the next 0.2-0.8 seconds, the high-frequency energy is relatively low and stable while the mid-frequency energy is very high; and in the last 0.8-1 seconds, the high-frequency energy suddenly increases while the mid-frequency energy returns to a stable state. Finally, these features are input into the language identification model, which determines that the language is English.

[0086] The beneficial effects of the above technologies are as follows: The spectrum generation submodule divides and transforms the speech signal by setting the frame length, frame shift, and moving window function to obtain the frequency domain spectrum, providing basic data for subsequent analysis. The resonance feature periodicity analysis submodule can analyze the resonance period and periodicity to represent the resonance change characteristics from the frequency domain spectrum, which helps to capture the characteristics of speech. The high-energy temporal distribution feature analysis submodule can analyze the high-energy temporal distribution characteristics, enriching the analysis dimensions of speech features. The interactive language recognition submodule inputs multiple features into the language recognition model, improving the accuracy and reliability of interactive language recognition. Overall, through the collaborative work of these submodules, the customer's interactive language can be more accurately identified, providing strong support for multilingual real-time translation POS terminal systems and improving the system's adaptability and service quality.

[0087] Example 3:

[0088] Based on Example 2, the resonance characteristic periodicity analysis submodule includes:

[0089] The resonance feature change trend curve generation unit is used to count the number of resonance peaks and the frequency points of resonance peaks in the spectrum of each frame of the first interactive speech data, and to perform time-series fitting on the number of resonance peaks and the frequency points of resonance peaks in the spectrum of all frames of the first interactive speech data to obtain the resonance peak change trend curve and the resonance peak frequency point change trend curve of the first interactive speech data.

[0090] The resonance period analysis unit is used to analyze the variation period of the formant variation trend curve and the variation period of the formant frequency point variation trend curve of the first interactive voice data based on the autocorrelation analysis method, and takes the least common multiple of the variation period of the formant variation trend curve and the variation period of the formant frequency point variation trend curve of the first interactive voice data as the resonance period of the first interactive voice data.

[0091] The periodic representative resonance variation feature analysis unit is used to input the formant variation trend curve, the formant frequency point variation trend curve, and the resonance period of the first interactive speech data into the periodic feature generalization model to obtain the periodic representative resonance variation features of the first interactive speech data.

[0092] In this embodiment, the number of formants and the frequency points of each frame of the initial interactive voice data are statistically determined as follows: The initial interactive voice data is processed to determine the number of formants and the frequency value corresponding to each formant in the spectrum of each frame of the voice data.

[0093] In this embodiment, the formant variation trend curve is a curve that reflects the change of the number of formants in each frame of the first interactive voice data by connecting them in chronological order.

[0094] In this embodiment, the formant frequency point change trend curve is a curve that connects the formant frequency points of each frame in the first interactive voice data in chronological order to show its change pattern.

[0095] In this embodiment, the variation period of the formant variation trend curve and the variation period of the formant frequency point variation trend curve of the first interactive voice data are analyzed based on the autocorrelation analysis method. The time interval corresponding to the repeated patterns in the formant variation trend curve and the formant frequency point variation trend curve is identified, i.e., the variation period.

[0096] In this embodiment, the periodic feature generalization model is a model capable of summarizing and refining the periodic features of speech data. For example, it is found that a frame of the first interactive speech data contains three formants, with frequencies of 500Hz, 1200Hz, and 2000Hz. Over time, the number of formants changes from 3 to 2 and back to 3, and the frequency of the formants fluctuates within the ranges of 500Hz-600Hz, 1200Hz-1300Hz, and 2000Hz-2100Hz. Through autocorrelation analysis, the period of the former is found to be 80 milliseconds, and the latter is found to be 100 milliseconds, thus determining the resonance period to be 400 milliseconds. Finally, these curves and the resonance period of 400 milliseconds are input into the periodic feature generalization model to obtain a more generalized description of the periodic features. Its construction process includes:

[0097] A large amount of speech data containing different languages, speakers, and contexts needs to be collected as training samples. These samples should cover all possible periodic features.

[0098] During the data preparation phase, the collected speech data is preprocessed, including operations such as framing, windowing, and spectrum analysis, in order to extract periodic features such as the formant variation trend curve and the formant frequency point variation trend curve.

[0099] Choose a suitable machine learning algorithm or deep learning architecture, such as a recurrent neural network (RNN) or a long short-term memory network (LSTM).

[0100] During model training, the extracted periodic features, the trend curves of formant peak changes, and the trend curves of formant peak frequency points are used as inputs, and the corresponding periodic feature summarization results are used as output labels. By continuously adjusting the model parameters, the model can learn the mapping relationship between the change period of different periodic change curves and the periodic feature summarization results.

[0101] The beneficial effects of the above technologies are as follows: The resonance feature change trend curve generation unit obtains the change trend curve by statistically analyzing and fitting the number of formants and frequency points, comprehensively presenting the changes in the resonance features of speech. The resonance period analysis unit accurately analyzes the change period of the formant and frequency point change trend curves using autocorrelation analysis, and determines the resonance period by calculating the least common multiple, thus improving the accuracy of period analysis. The periodic representative resonance change feature analysis unit inputs the correlation curve and resonance period into the periodic feature generalization model to obtain more representative periodic resonance change features, which helps to improve the accuracy of subsequent language recognition. Overall, the collaborative work of these units enables a deeper and more accurate analysis of the resonance features of speech, providing strong support for the accurate recognition of interactive languages ​​and enhancing the performance and adaptability of the multilingual real-time translation POS terminal system.

[0102] Example 4:

[0103] Based on Example 2, the high-energy time-domain distribution characteristic analysis submodule includes:

[0104] The frequency spectrum centroid change trend curve fitting unit is used to determine the frequency spectrum centroid of each frame of the speech signal based on the frequency domain spectrum of each frame of the first interactive speech data, and to perform time-series fitting on the frequency spectrum centroid of all frames of the first interactive speech data to obtain the frequency spectrum centroid change trend curve of the first interactive speech data.

[0105] The energy concentration frequency band change trend curve fitting unit is used to simultaneously divide the frequency band and analyze the frequency band energy distribution characteristics of the frequency domain spectrum of each frame of the first interactive voice data to obtain the energy concentration frequency band change trend curve of the first interactive voice data.

[0106] The merged feature analysis unit is used to input the spectral centroid change trend curve and the energy concentration frequency band change trend curve of the first interactive speech data into the merged feature analysis model to obtain the high-energy time-domain distribution characteristics of the first interactive speech data.

[0107] In this embodiment, the spectral centroid of each frame of speech signal is determined based on the frequency domain spectrum of each frame of speech signal in the initial interactive speech data. This is done by calculating the frequency value representing the center of energy distribution for that frame, i.e., the spectral centroid, according to the frequency domain distribution of each frame of speech signal in the initial interactive speech data. For example, assuming that the low-frequency part of a frame of speech signal has stronger energy and the high-frequency part has weaker energy, its spectral centroid is calculated to be 500Hz.

[0108] In this embodiment, the spectral centroid change trend curve of the initial interactive voice data is formed by connecting the spectral centroids of each frame of the voice signal in the initial interactive voice data in chronological order, reflecting the change of the spectral centroid over time. For example, in the first half of a voice segment, the spectral centroid gradually rises from 300Hz to 600Hz, and then drops to 400Hz in the second half. Connecting these points forms the spectral centroid change trend curve.

[0109] In this embodiment, the energy concentration frequency band change trend curve represents the curve of the frequency band with concentrated energy in the initial interactive voice data changing over time. For example, in a certain voice segment, the energy is initially concentrated in the 1000-2000Hz frequency band, and as time goes on, the energy concentration frequency band changes to 500-1000Hz. This change is represented by a curve.

[0110] In this embodiment, the merged feature analysis model is used to comprehensively analyze and process multiple related features (such as spectral centroid and energy concentration bands). This model can receive inputs such as spectral centroid change trend curves and energy concentration band change trend curves, and then output more comprehensive and in-depth analysis results of speech features. Its construction process includes:

[0111] It is necessary to collect a rich variety of speech data as training samples. These samples should have different characteristics such as spectral centroid change trend curves and energy concentration frequency band change trend curves.

[0112] The speech data is preprocessed to obtain features such as the spectral centroid and energy concentration bands, and their changing trends are calculated.

[0113] Choose a model structure suitable for handling multi-feature fusion, such as a combination of multilayer perceptron (MLP) or convolutional neural network (CNN) and RNN / LSTM.

[0114] During training, features such as the trend curve of spectral centroid change and the trend curve of energy concentration frequency band change are simultaneously input into the model. The desired output is a comprehensive analysis result of these features. By repeatedly optimizing the model parameters, it is made possible to accurately combine and analyze multiple input features.

[0115] During the training of these two models, the training samples should be as diverse as possible, including speech under conditions such as different languages, speech rates, intonations, accents, and noisy environments, to ensure that the models have good generalization ability and robustness.

[0116] The beneficial effects of the above technologies are as follows: The spectral centroid change trend curve fitting unit obtains the spectral centroid change trend curve by determining and fitting the spectral centroid, which reflects the changes in the energy distribution of the speech signal. The energy concentration frequency band change trend curve fitting unit divides the frequency band and analyzes the energy distribution characteristics to obtain the energy concentration frequency band change trend curve, describing the energy characteristics of speech from different perspectives. The merging feature analysis unit inputs the spectral centroid change trend curve and the energy concentration frequency band change trend curve into the merging feature analysis model to obtain the high-energy temporal distribution characteristics, comprehensively considering multiple energy distribution information. Overall, the collaborative work of these units can more comprehensively and accurately analyze the high-energy temporal distribution characteristics of the initial interactive speech data, providing richer and more accurate feature information for subsequent interactive language recognition, and helping to improve the accuracy and efficiency of language recognition in multilingual real-time translation POS terminal systems.

[0117] Example 5:

[0118] Based on Example 4, the energy concentration frequency band variation trend curve fitting unit includes:

[0119] The frequency band division and frequency band energy calculation subunit is used to simultaneously divide the frequency domain spectrum of each frame of the first interactive voice data into frequency bands, obtain multiple frequency bands of each frame of the first interactive voice data, and calculate the frequency band energy of each frequency band of each frame of the voice data.

[0120] The normalization and curve fitting subunit is used to normalize the frequency band energy of all the same frequency band in all frames of the first interactive voice data and then fit it according to the time sequence to obtain the energy change trend curve of each frequency band of the first interactive voice data.

[0121] The curve alignment analysis subunit is used to perform alignment analysis on the energy change trend curves of all frequency bands of the first interactive voice data, and to fit the energy concentration frequency band change trend curve of the first interactive voice data.

[0122] In this embodiment, the frequency domain spectrum of each frame of the initial interactive voice data is simultaneously divided into frequency bands: the spectral range of each frame of the initial interactive voice data in the frequency domain is divided into several frequency bands. For example, the 0-3000Hz frequency range can be divided into frequency bands such as 0-1000Hz, 1000-2000Hz, and 2000-3000Hz.

[0123] In this embodiment, the frequency band energy of each frequency band in each frame of the speech signal is calculated as follows: for each divided frequency band of each frame of the speech signal, the total energy within that frequency band is calculated. For example, the energy value within the 0-1000Hz frequency band in a certain frame is 10 joules.

[0124] In this embodiment, the frequency band energy of all frames of the initial interactive voice data in the same frequency band is normalized and then fitted according to time sequence to obtain the energy change trend curve of each frequency band of the initial interactive voice data. This is achieved by processing the energy of the same frequency band in all frames to a unified standard (normalization), and then connecting these energy values ​​in time sequence to form a curve reflecting the change trend of frequency band energy over time. For example, assuming the normalized energy of the 0-1000Hz frequency band is 0.2, 0.3, and 0.4 in the first few frames, an upward curve is fitted according to time sequence.

[0125] In this embodiment, the energy change trend curve for each frequency band refers to the curve obtained above for each frequency band reflecting the change of its energy over time. For example, the 0-1000Hz frequency band has an upward energy change trend curve, and the 1000-2000Hz frequency band has a curve that first decreases and then increases.

[0126] In this embodiment, the energy change trend curves of all frequency bands of the initial interactive voice data are aligned and analyzed to fit the energy concentration frequency band change trend curve of the initial interactive voice data. The energy change trend curves of each frequency band are compared and analyzed together to find the frequency bands with relatively concentrated energy at each time sequence, and all frequency bands with relatively concentrated energy at all time sequences are fitted into a curve (the horizontal axis is the time sequence, and the vertical axis is the frequency band range value). For example, through analysis, it is found that the energy of the 1000-2000Hz frequency band is relatively concentrated from 0 seconds to 0.5 seconds, while the energy of the 2000-3000Hz frequency band is relatively concentrated from 0.5 seconds to 1.2 seconds, and the energy of the 0-1000Hz frequency band is relatively concentrated from 1.2 seconds to 3 seconds. Therefore, their frequency band range values ​​are connected according to the time sequence to fit a new curve, that is, the energy concentration frequency band change trend curve.

[0127] The beneficial effects of the above technologies are as follows: The frequency band division and frequency band energy calculation subunit performs frequency band division and energy calculation on the speech signal, providing detailed frequency band energy information. The normalization and curve fitting subunit obtains the energy change trend curve of each frequency band through normalization and fitting, making the energy changes of different frequency bands comparable. The curve alignment analysis subunit performs alignment analysis on the energy change trend curves of all frequency bands, fitting the frequency band change trend curve with concentrated energy, which more accurately reflects the frequency band changes with concentrated energy. Overall, the collaborative work of these subunits can accurately analyze the frequency band energy distribution characteristics of the speech signal, providing reliable data support for obtaining high-energy time-domain distribution characteristics, and helping to improve the ability of multilingual real-time translation POS terminal systems to analyze interactive speech data and the accuracy of language recognition.

[0128] Example 6:

[0129] Based on Example 1, the translation and ambiguous content recognition module, referring to Figure 3, includes:

[0130] The real-time customer voice translation submodule is used to convert the real-time interactive voice data into text by treating the customer's interactive language as the input language when it receives interactive voice data in real time, so as to obtain the customer's voice input text.

[0131] The customer ambiguous word retrieval submodule is used to take the cashier's interactive language as the output language and retrieve all the definitions of each word in the customer's voice input text in the output language from the multilingual lexicon. It also takes words in the customer's voice input text that contain more than one definition in the output language as ambiguous words in the corresponding interactive voice data.

[0132] The Customer Ambiguous Sentence Retrieval Submodule is used to perform grammatical deconstruction on each sentence in the customer's voice input text to obtain all the grammatical structures of each sentence, and to treat sentences in the customer's voice input text that contain more than one grammatical structure as ambiguous sentences in the corresponding interactive voice data.

[0133] The Customer Alternate Translation Result Generation Submodule is used to generate all alternative translation results for the corresponding interactive voice data in the corresponding output language, based on all ambiguous words and sentences in the interactive voice data input by the customer in real time, as well as the customer's voice input text.

[0134] In this embodiment, customer voice input text refers to the text content obtained after the customer's real-time interactive voice input has been converted. For example, if the customer says "I want to buy this apple," the converted text will be "I want to buy this apple."

[0135] In this embodiment, the multilingual semantic database is a database that stores words in multiple languages ​​and their definitions. For example, the English word "charge" has multiple definitions such as "to charge; to accuse; to manage; to recharge".

[0136] In this embodiment, performing grammatical deconstruction on each sentence in the customer's voice input text to obtain all possible grammatical structures means analyzing the sentence components and grammatical rules of each sentence in the customer's input text to find all possible grammatical configurations. For example, the sentence "Isawamanwithatelescope." can be understood grammatically as "I saw a man with a telescope," where "aman" is the object being seen and "withatelescope" modifies "man." However, it could also be understood as "I saw a man with a telescope," where "I" saw "man" through the tool "telescope."

[0137] In this embodiment, based on all ambiguous words and sentences in the real-time interactive voice data input by the customer, as well as the customer's voice input text, all alternative translation results for the corresponding interactive voice data in the corresponding output language are generated as follows:

[0138] Based on the multiple-meaning words and ambiguous sentences in the customer's real-time speech, as well as the entire input text, multiple possible translations of this speech in the output language are generated. For example, if the customer says "Visiting relative can be boring," "visiting" is an ambiguous word. One interpretation is "Visiting relative can be boring," with "Visiting relative" as a gerund phrase as the subject; another interpretation is "Visiting relative can be boring," where "Visiting" is an adjective modifying "relatives." Therefore, both of these alternative translations can be generated.

[0139] The beneficial effects of the above technologies are as follows: The real-time customer voice translation submodule can promptly convert customer interactive voice data into text, providing a foundation for subsequent analysis and processing. The customer ambiguous word retrieval submodule accurately identifies ambiguous words in the customer's voice input text by searching in a multilingual semantic database, which helps to comprehensively identify ambiguities in translation. The customer ambiguous sentence retrieval submodule performs grammatical deconstruction on sentences to identify ambiguous sentences containing multiple grammatical structures, further improving the identification of ambiguous content. The customer alternative translation result generation submodule generates all alternative translation results based on ambiguous words, ambiguous sentences, and voice input text, providing more possibilities for subsequent selection of accurate translations. Overall, these modules work together to effectively translate customer-input interactive voice data and identify ambiguous content, improving the accuracy and comprehensiveness of translation, thereby enhancing the service quality and efficiency of the multilingual real-time translation POS terminal system.

[0140] Example 7:

[0141] Based on Example 1, the translation and ambiguous content recognition module, referring to Figure 4, includes:

[0142] The cashier's real-time voice translation submodule is used to convert the real-time interactive voice data into text when the cashier's interactive language is received in real time, so as to obtain the cashier's voice input text.

[0143] The cashier ambiguous word retrieval submodule is used to take the customer's interaction language as the output language and retrieve all the definitions of each word in the cashier's voice input text in the output language from the multilingual lexicon. Words in the cashier's voice input text that contain more than one definition in the output language are treated as ambiguous words in the corresponding interactive voice data.

[0144] The Cashier Ambiguity Retrieval Submodule is used to perform grammatical deconstruction on each sentence in the cashier's voice input text to obtain all the grammatical structures of each sentence, and to treat sentences in the cashier's voice input text that contain more than one grammatical structure as ambiguous sentences in the corresponding interactive voice data.

[0145] The cashier alternative translation result generation submodule is used to generate all alternative translation results for the corresponding interactive voice data in the corresponding output language based on all ambiguous words and sentences in the interactive voice data input by the cashier in real time, as well as the cashier's voice input text.

[0146] In this embodiment, the cashier's voice input text refers to the text content obtained by converting the interactive voice input by the cashier in real time. For example, if the cashier says "This item is 20% off," the converted text will be "This item is 20% off."

[0147] In this embodiment, the grammatical structure of each sentence in the cashier's voice input text is deconstructed to obtain all the grammatical structures of each sentence: that is, each sentence in the text input by the cashier is analyzed from a grammatical perspective to find out all possible composition methods and rules.

[0148] In this embodiment, based on all ambiguous words and sentences in the real-time interactive voice data input by the cashier, as well as the cashier's voice input text, all alternative translation results for the corresponding interactive voice data in the corresponding output language are generated: based on words with multiple interpretations, ambiguous sentences, and the entire input text in the cashier's real-time speech, all possible translation results for this speech in the corresponding output language are obtained.

[0149] The beneficial effects of the above technologies are as follows: The real-time voice translation submodule for cashiers can promptly convert the cashier's interactive voice data into text, facilitating subsequent processing. The cashier ambiguous word retrieval submodule searches the multilingual semantic database to identify ambiguous words in the cashier's input text, improving the comprehensive ability to identify ambiguities. The cashier ambiguous sentence retrieval submodule identifies ambiguous sentences through grammatical deconstruction, improving the recognition of ambiguous content in the cashier's input voice. The cashier alternative translation result generation submodule generates alternative translation results based on ambiguous words, ambiguous sentences, and input text, providing more options for selecting accurate translations. Overall, these modules work together to effectively translate the cashier's input interactive voice data and identify ambiguous content, improving the accuracy and completeness of the translation, and enhancing the reliability and service effectiveness of the multilingual real-time translation POS terminal system in practical applications.

[0150] Example 8:

[0151] Based on Example 1, the alternative translation filtering module, referring to Figure 5, includes:

[0152] The word meaning filtering submodule is used to determine the actual meaning of each ambiguous word based on the nearby context of each ambiguous word in the real-time input interactive voice data and the preset multi-word meaning weight information database of the cashier scenario.

[0153] The grammatical structure filtering submodule is used to determine the weight value of each grammatical structure of each ambiguous sentence based on the neighboring context of each ambiguous sentence, and to take the grammatical structure with the largest weight value among all grammatical structures of each ambiguous sentence as the actual grammatical structure of each ambiguous sentence.

[0154] The translation result filtering submodule is used to filter the final translation result of the real-time input interactive voice data from all candidate translation results based on the actual meaning of all ambiguous words and the actual grammatical structure of all ambiguous sentences in the real-time input interactive voice data.

[0155] In this embodiment, the actual meaning of an ambiguous word is its true meaning in a specific context. For example, in the context of "went to the bank," the actual meaning of "bank" is "bank."

[0156] In this embodiment, determining the weight value of each grammatical structure of each ambiguous sentence based on its surrounding context means assigning a weight value to each possible grammatical structure of the ambiguous sentence based on the content of related statements around it. For example, suppose the ambiguous sentence is: "I found the book on the ground." It has two possible grammatical structures:

[0157] 1. It can be understood as "I found (the book on the ground)," where "the book on the ground" is an object clause indicating the specific content of the discovery;

[0158] 2. It can be understood as "I found the book (fallen on the ground)". "Fallen on the ground" is the complement of "book" and describes the state of the book.

[0159] The adjacent context is: "I quickly picked it up and put it back on the bookshelf." For the first grammatical structure, since the subsequent context mainly revolves around the subsequent action of "the book falling on the ground," it is given a higher initial weight, such as 0.7. For the second grammatical structure, since the subsequent context does not directly describe or emphasize the state of the "book," it is given a lower initial weight, such as 0.3.

[0160] One possible way to calculate the weights is to start with an initial weight for each grammatical structure and then adjust it based on factors such as keywords, semantic connections, and logical coherence related to each structure in the surrounding context. If more words and expressions in the surrounding context fit the logic of the first grammatical structure better, its weight is increased, and the weight of the second grammatical structure is decreased. The adjustment range of the weights can be determined based on the degree of fit; for example, a high degree of fit increases the weight by 0.2, a low degree of fit increases it by 0.1, and so on. This ultimately yields the final weighted value for each grammatical structure.

[0161] In this embodiment, the actual grammatical structure of an ambiguous sentence refers to the grammatical structure that is ultimately determined to be used in a specific context.

[0162] In this embodiment, based on the actual meaning of all ambiguous words and the actual grammatical structure of all ambiguous sentences in the real-time input interactive voice data, the final translation result of the real-time input interactive voice data is selected from all corresponding candidate translation results. This means that according to the definite meaning of ambiguous words and the definite grammatical structure of ambiguous sentences in the interactive voice, the most suitable and accurate translation is chosen as the final translation result from the previously obtained multiple candidate translations. For example, by analyzing and determining that a certain ambiguous word has meaning A and a certain ambiguous sentence has grammatical structure B, the final translation that meets these conditions is selected from multiple candidates.

[0163] The beneficial effects of the above technologies are as follows: The semantic filtering submodule accurately determines the actual meaning of each ambiguous word by using its proximity context and a multi-semantic weight information database of the checkout scenario, thus improving the accuracy of semantic selection. The grammatical structure filtering submodule determines the weight value of the grammatical structure based on the proximity context of the ambiguous sentence, selecting the actual grammatical structure and enhancing the reliability of understanding the ambiguous sentence. The translation result filtering submodule filters the final translation result based on the actual meaning of the ambiguous word and the actual grammatical structure of the ambiguous sentence, ensuring the accuracy and rationality of the translation. Overall, these submodules work together to effectively select the final translation result that best matches the actual context from the candidate translation results, improving the translation quality and practicality of the multilingual real-time translation checkout terminal system, reducing communication barriers caused by ambiguity, and improving checkout efficiency and service satisfaction.

[0164] Example 9:

[0165] Based on Example 8, the word meaning filtering submodule includes:

[0166] The total weighting unit is used to determine the total weighting value of each meaning of each ambiguous word based on the nearby context of each ambiguous word in the real-time input interactive voice data and the preset multi-meaning weight information database of the cashier scenario.

[0167] The definition filtering unit is used to take the definition with the highest total weight among all the definitions of each ambiguous word as the actual meaning of each ambiguous word.

[0168] In this embodiment, the total weight value of the definition refers to a comprehensive weight value calculated for a given definition of an ambiguous word, taking into account its contextual situation and the weights in a pre-defined multi-semantic weight information database for checkout scenarios. For example, the ambiguous word "book" has two definitions: "book" and "reservation." In a specific context, by analyzing the surrounding context and referring to the multi-semantic weight information database for checkout scenarios, the total weight value for the definition "book" is calculated to be 0.2, and the total weight value for the definition "reservation" is 0.8. The total weight value is used to determine the actual meaning of the ambiguous word in that context.

[0169] The beneficial effects of the above technologies are as follows: The total weighting unit determines the total weight value for the definition of each ambiguous word by using the adjacent context of the ambiguous word and the multi-semantic weight information database of the checkout scenario, making the evaluation of word meaning more quantitative and objective. The definition screening unit takes the definition with the highest total weight value as the actual meaning, ensuring the accuracy and rationality of the selection. Overall, the collaborative work of these two units can accurately determine the actual meaning of ambiguous words, improve the screening accuracy of the alternative translation screening module, and thus further enhance the translation accuracy and practicality of the multilingual real-time translation checkout terminal system, providing more reliable support for cross-language communication, reducing misunderstandings caused by semantic ambiguity, and optimizing the communication effect and service quality during the checkout process.

[0170] Example 10:

[0171] Based on Example 9, the total weighting unit includes:

[0172] The first weighting subunit is used to determine the first weighting value for each interpretation of each ambiguous word based on the neighborhood context of each ambiguous word in real-time input interactive voice data.

[0173] The second weighting subunit is used to retrieve the preset multi-semantic weight information database for the cashier scenario and determine the second weighting value for each definition of each ambiguous word.

[0174] The total weighting subunit is used to take the sum of the first and second weighting values ​​of each definition of each ambiguous word as the total weighting value of each definition.

[0175] In this embodiment, determining the first weighting value for each ambiguous word's meaning based on its neighboring context in real-time input interactive voice data refers to assigning an initial weight value to each possible meaning of the ambiguous word based on the surrounding linguistic environment of the ambiguous word in the real-time received interactive voice. For example, the ambiguous word "light" appears in the interactive voice, which can mean "light" or "light." Suppose the sentence is: "The light in the room is too dim." In this sentence, "light" is more clearly defined as "light," so its initial weight is higher. Let's assume we set the initial weight for "light" to 0.8 and the initial weight for "light" to 0.2 (the initial weights can be determined by retrieving a pre-set information database containing the initial weights corresponding to different sentence meanings). Then, we consider related sentences, such as the following sentence "We need a brighter light bulb." This further strengthens the possibility that "light" means "light." A simple weighted average method can be used to adjust the weight values. Assuming the enhancement weight is 0.2, the new weight is calculated as follows:

[0176] The new weight of "light" (i.e., the first weighting value) = 0.8 + 0.2 × 0.8 = 0.96;

[0177] The new weight for "light" (i.e., the first weight) = 0.2 - 0.2 × 0.2 = 0.16.

[0178] The beneficial effects of the above technologies are as follows: The first weighting subunit determines the first weighting value based on the adjacent context of the ambiguous word, fully considering the contextual information of the current communication, making the weighting more consistent with the actual context. The second weighting subunit determines the second weighting value by retrieving the multi-semantic weight information database of the checkout scenario, introducing semantic weights under specific scenarios, increasing the professionalism and targeting of the weighting. The total weighting subunit adds the two weighting values ​​to obtain the total weighting value, integrating contextual and professional scenario factors, making the total weighting more comprehensive and accurate. Overall, the collaborative work of these subunits can scientifically and accurately determine the total weighting value for each interpretation of the ambiguous word, providing a reliable basis for accurately filtering the actual meaning of the ambiguous word, thereby improving the performance of the alternative translation screening module, further improving the translation accuracy and adaptability of the multilingual real-time translation checkout terminal system in complex language environments, and providing a higher quality and more efficient service for cross-language checkout communication.

[0179] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A multilingual real-time translation POS terminal system, characterized in that, include: The interactive language recognition module is used to determine the customer's interactive language based on the initial interactive voice data input by the customer; The translation and ambiguous content recognition module is used to translate the interactive voice data input in real time by customers or cashiers based on the customer's or cashier's interactive language, and to identify all ambiguous words, all ambiguous sentences, and all alternative translation results in the corresponding interactive voice data. The alternative translation filtering module is used to filter the final translation result of the real-time input interactive voice data from all candidate translation results, based on the neighboring contexts of all ambiguous words and sentences in the real-time input interactive voice data, as well as a preset multi-semantic weight information database for the cashier scenario. The interactive language recognition module includes a spectrum generation submodule, which is used to divide the voice signal of the initial interactive voice data input by the customer based on a preset frame length, frame shift, and moving window function, obtaining the windowed voice signal of each frame of the initial interactive voice data, and performing a Fast Fourier Transform on the windowed voice signal of each frame of the initial interactive voice data to obtain the initial interactive voice... The system comprises: a frequency domain spectrum analysis module for each frame of the initial interactive speech data; a periodicity analysis module for resonant features, used to analyze the frequency domain spectrum of each frame of the initial interactive speech data to determine the resonant period and periodic representative resonant variation characteristics; a high-energy time-domain distribution feature analysis module for analyzing the high-energy time-domain distribution characteristics of each frame of the initial interactive speech data; and an interactive language recognition module, used to input the resonant period, periodic representative resonant variation characteristics, and high-energy time-domain distribution characteristics of the initial interactive speech data into the language recognition model to obtain the customer's interactive language. The periodicity analysis module for resonant features... The system includes: a resonance feature trend curve generation unit, used to statistically determine the number of formants and the frequency points of each frame of the initial interactive speech data in the spectrogram, and to perform time-series fitting on the number of formants and the frequency points of all frames of the initial interactive speech data to obtain the formant trend curve and the frequency point trend curve; and a resonance period analysis unit, used to analyze the period of the formant trend curve and the frequency point trend curve of the initial interactive speech data based on autocorrelation analysis, and to combine the period of the formant trend curve and the frequency point trend curve of the initial interactive speech data. The least common multiple of the variation cycles of the formant frequency point change trend curve is taken as the resonance cycle of the first interactive voice data; the periodic representative resonance change feature analysis unit is used to input the formant change trend curve, the formant frequency point change trend curve, and the resonance cycle of the first interactive voice data into the periodic feature generalization model to obtain the periodic representative resonance change features of the first interactive voice data; among them, the translation and ambiguous content recognition module includes: a real-time customer voice translation submodule, which is used to treat the customer's interactive language as the input language when real-time interactive voice data is received from the customer to perform text conversion on the real-time input interactive voice data to obtain the customer's voice input text;The customer ambiguous word retrieval submodule takes the cashier's interactive language as the output language and retrieves all definitions of each word in the customer's voice input text in the output language from a multilingual lexical database. Words in the customer's voice input text that contain more than one definition in the output language are treated as ambiguous words in the corresponding interactive voice data. The customer ambiguous sentence retrieval submodule performs grammatical deconstruction on each sentence in the customer's voice input text to obtain all grammatical structures for each sentence. Sentences in the customer's voice input text that contain more than one grammatical structure are treated as ambiguous sentences in the corresponding interactive voice data. The customer alternative translation result generation submodule generates corresponding interactive voice data based on all ambiguous words and sentences in the real-time customer-input interactive voice data and the customer's voice input text. The system should output all candidate translation results for the specified language. The candidate translation filtering module includes: a semantic filtering submodule, used to determine the actual meaning of each ambiguous word based on its neighboring context and a pre-defined multi-semantic weight information database for the cashier scenario; a grammatical structure filtering submodule, used to determine the weight value of each grammatical structure of each ambiguous sentence based on its neighboring context, and to use the grammatical structure with the highest weight among all grammatical structures of each ambiguous sentence as its actual grammatical structure; and a translation result filtering submodule, used to filter the final translation result of the real-time input interactive voice data from all candidate translation results based on the actual meaning of all ambiguous words and the actual grammatical structure of all ambiguous sentences in the real-time input interactive voice data.

2. The multilingual real-time translation POS terminal system according to claim 1, characterized in that, The high-energy temporal distribution characteristic analysis submodule includes: a spectral centroid change trend curve fitting unit, used to determine the spectral centroid of each frame of the speech signal based on the frequency domain spectrum of each frame of the initial interactive speech data, and to perform time-series fitting on the spectral centroids of all frames of the initial interactive speech data to obtain the spectral centroid change trend curve of the initial interactive speech data; an energy concentration frequency band change trend curve fitting unit, used to simultaneously perform frequency band division and frequency band energy distribution characteristic analysis on the frequency domain spectrum of each frame of the initial interactive speech data to obtain the energy concentration frequency band change trend curve of the initial interactive speech data; and a merging feature analysis unit, used to input the spectral centroid change trend curve and the energy concentration frequency band change trend curve of the initial interactive speech data into the merging feature analysis model to obtain the high-energy temporal distribution characteristics of the initial interactive speech data.

3. The multilingual real-time translation POS terminal system according to claim 2, characterized in that, The energy concentration frequency band change trend curve fitting unit includes: a frequency band division and frequency band energy calculation subunit, which is used to simultaneously divide the frequency domain spectrum of each frame of the first interactive speech data into frequency bands, obtain multiple frequency bands of each frame of the first interactive speech data, and calculate the frequency band energy of each frequency band of each frame of the speech data; a normalization and curve fitting subunit, which is used to normalize the frequency band energy of all the same frequency bands of all frames of the first interactive speech data and then fit them according to the time sequence to obtain the energy change trend curve of each frequency band of the first interactive speech data; and a curve alignment analysis subunit, which is used to perform alignment analysis on the energy change trend curves of all frequency bands of the first interactive speech data and fit the energy concentration frequency band change trend curve of the first interactive speech data.

4. The multilingual real-time translation POS terminal system according to claim 1, characterized in that, The translation and ambiguous content recognition module includes: a real-time cashier voice translation submodule, which, when receiving real-time interactive voice data input from a cashier, treats the cashier's interactive language as the input language and performs text conversion on the real-time input interactive voice data to obtain the cashier's voice input text; and a cashier ambiguous word retrieval submodule, which treats the customer's interactive language as the output language and retrieves all definitions of each word in the cashier's voice input text in the output language from a multilingual semantic database, identifying words in the cashier's voice input text that contain more than one definition in the output language. The first module treats ambiguous words as words in the corresponding interactive voice data; the second module retrieves ambiguous sentences from the cashier's voice input text by performing grammatical deconstruction on each sentence to obtain all grammatical structures of each sentence, and treats sentences containing more than one grammatical structure in the cashier's voice input text as ambiguous sentences in the corresponding interactive voice data; the third module generates alternative translation results for the cashier's voice input text by generating all alternative translation results in the corresponding output language based on all ambiguous words and sentences in the cashier's real-time input interactive voice data and the cashier's voice input text.

5. The multilingual real-time translation POS terminal system according to claim 1, characterized in that, The word meaning filtering submodule includes: a total weighting unit, which determines the total weight value of each definition of each ambiguous word based on the nearby context of each ambiguous word in the real-time input interactive voice data and the preset multi-word meaning weight information database for the cashier scenario; and a definition filtering unit, which takes the definition with the largest total weight value among all the definitions of each ambiguous word as the actual meaning of each ambiguous word.

6. The multilingual real-time translation POS terminal system according to claim 5, characterized in that, The total weighting unit includes: a first weighting subunit, used to determine a first weighting value for each definition of each ambiguous word based on the neighboring context of each ambiguous word in real-time input interactive voice data; a second weighting subunit, used to retrieve a preset multi-semantic weight information database for cashier scenarios and determine a second weighting value for each definition of each ambiguous word; and a total weighting subunit, used to take the sum of the first weighting value and the second weighting value for each definition of each ambiguous word as the total weighting value for each definition.

Citation Information

Patent Citations

  • Cash register and cashier system with translation function

    CN204044940U

  • Method for selecting target word in dialogue automatic translation

    KR1020130022473A