A Multilingual Concurrent Recognition Method and System Based on AI
By using AI neural network models and time-division multiplexing scheduling technology, multilingual audio signals are encoded and decoded, solving the problem of concurrent recognition and transmission of multiple languages within a single channel, and achieving efficient and accurate multilingual communication.
Patent Information
- Application Number
- CN202410563537.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2044-05-08
AI Technical Summary
Existing AI translation technologies struggle to efficiently recognize and transmit multiple languages concurrently within a single channel, especially in high-frequency communication environments where accuracy is low and confusion is common.
An AI neural network model is used to train multilingual audio signals to determine the language recognition information set. The language recognition information is then processed through frequency coding, wavelength characteristic coding, or a combination of coding. Combined with a time-division multiplexing scheduling plan, the coded signal is scheduled into a composite multi-track signal, which is then transmitted and decoded into a recognizable target multilingual audio signal in a single channel.
It enables efficient concurrent recognition and transmission of multiple languages within a single channel, improving communication efficiency and accuracy while avoiding interference and confusion between signals.
Smart Images

Figure CN118335063B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent language communication, and in particular to an AI-based multi-language concurrent recognition method and system. BACKGROUND
[0002] With the vigorous development of the global economy and the deepening of cultural exchanges, people's demand for friendly communication among multiple languages is increasing. Artificial intelligence translation technology has gradually become popular, which can automatically recognize and translate unfamiliar languages, making cross-cultural communication more convenient and efficient, and providing important support for business, tourism, academic exchanges and other fields.
[0003] However, existing artificial intelligence translation technology mainly uses a single channel to translate a single dialogue or text, that is, it can only process the recognition and translation of one language at the same time, and it is difficult to efficiently recognize and transmit multiple languages in a single channel, especially in a high-frequency communication environment, the multi-language recognition accuracy is low, and it is easy to cause confusion, and it is difficult to support simultaneous processing and transmission of a large number of languages in the case of limited single-channel resources.
[0004] The existing artificial intelligence translation technology cannot concurrently recognize and transmit multiple languages in a single channel, which limits the communication efficiency in a multi-language environment. SUMMARY
[0005] Therefore, the present application provides an AI-based multi-language concurrent recognition method and system to solve the technical problem that artificial intelligence translation technology cannot concurrently recognize and transmit multiple languages in a single channel.
[0006] In a first aspect, the present application provides an AI-based multi-language concurrent recognition method, which adopts the following technical solution:
[0007] Obtain a multi-language audio signal;
[0008] Train the multi-language audio signal according to an AI neural network model to obtain a language recognition information set;
[0009] Determine a target encoding mode according to the language recognition information set;
[0010] Encode the language recognition information set according to the target encoding mode to obtain a target encoded signal;
[0011] Determine a time division multiplexing scheduling plan according to the target encoded signal;
[0012] Schedule the target encoded signal into a composite multi-audio track signal according to the time division multiplexing scheduling plan;
[0013] According to the target coding mode and the time division multiplexing scheduling plan, the composite multi-track signal is decoded to obtain a target multi-language audio signal.
[0014] In an optional implementation, the determining of the target coding mode according to the set of language recognition information includes:
[0015] The language type and language start and end time of each language recognition information in the set of language recognition information are obtained.
[0016] According to the language type, data matching is performed in an encoding strategy database to determine the encoding mode of each language recognition information, and the encoding mode includes frequency encoding, wavelength characteristic encoding, and combined encoding.
[0017] According to the language start and end time, the time sequence position of the encoding mode is determined.
[0018] According to the time sequence position, the encoding mode is sorted to obtain a target coding mode.
[0019] In an optional implementation, the encoding of the set of language recognition information according to the target coding mode to obtain a target encoding signal includes:
[0020] A first set of recognition information in the set of language recognition information is obtained, and the following operations are performed on the first set of recognition information:
[0021] Step A1, a preset frequency band is obtained, the first set of recognition information is frequency band allocated based on the preset frequency band, and a target frequency band set is determined.
[0022] Step A2, the first set of recognition information is mapped to the target frequency band set to obtain a first set of encoding signals.
[0023] A second set of recognition information in the set of language recognition information is obtained, and the following operations are performed on the second set of recognition information:
[0024] Step B1, a preset wavelength band is obtained, and the second set of recognition information is wavelength allocated to determine a target wavelength set.
[0025] Step B1, a preset wavelength band is obtained, and the second set of recognition information is wavelength allocated based on the preset wavelength band to determine a target wavelength band set.
[0026] Step B2, the second set of recognition information is mapped to the target wavelength band set to obtain a second set of encoding information.
[0027] acquiring a third set of identification information from the set of language identification information, the encoding mode of which is the set encoding, and performing the following operations on the third set of identification information:
[0028] Step C1, acquiring a frequency encoding area and a wavelength characteristic encoding area in the third set of identification information;
[0029] Step C2, performing the frequency encoding on the frequency encoding area and performing the wavelength characteristic encoding on the wavelength characteristic encoding area to obtain a third set of encoded information;
[0030] According to the first set of encoded signals, the second set of encoded signals, and the third set of encoded signals, the target encoded signal is obtained by combination.
[0031] In an optional embodiment, the determining of the time division multiplexing scheduling plan according to the target encoded signal comprises:
[0032] According to the language start and end time, the language identification information time slice of each language identification information is determined;
[0033] According to the language type, the language identification information priority of each language identification information is determined;
[0034] According to the language identification information time slice and the language identification information priority, the time division multiplexing scheduling plan is determined.
[0035] In an optional embodiment, the scheduling of the target encoded signal into a composite multi-audio track signal according to the time division multiplexing scheduling plan comprises:
[0036] According to the language identification information time slice, the time slice window of the target encoded signal is determined;
[0037] According to the language identification information priority, the time slice window priority is determined;
[0038] According to the time slice window priority, the scheduling is performed to obtain the composite multi-audio track signal.
[0039] In an optional embodiment, the decoding processing of the composite multi-audio track signal according to the target encoding mode and the time division multiplexing scheduling plan to obtain an identifiable target multi-language audio signal comprises:
[0040] According to the time division multiplexing scheduling plan, the timing decoding is performed to obtain first decoding information;
[0041] The frequency decoding is performed on the first decoding information to obtain second decoding information;
[0042] performing wavelength decoding on the second decoding information to obtain third decoding information;
[0043] performing multi-language audio reconstruction on the third decoding information to obtain the target multi-language audio signal.
[0044] In an optional implementation, the method further includes:
[0045] monitoring a channel state in real time, and adjusting the target encoding mode and the time division multiplexing scheduling plan according to the channel state.
[0046] In a second aspect, the present application provides an AI-based multi-language concurrent recognition system, which includes:
[0047] a signal acquisition module, configured to acquire a multi-language audio signal;
[0048] a language recognition module, configured to train the multi-language audio signal according to an AI neural network model to obtain a language recognition information set;
[0049] a signal encoding module, configured to determine a target encoding mode according to the language recognition information set;
[0050] The signal encoding module is further configured to encode the language recognition information set according to the target encoding mode to obtain a target encoding signal.
[0051] a multiplexing scheduling module, configured to determine a time division multiplexing scheduling plan according to the target encoding signal;
[0052] a track synthesizing module, further configured to schedule the target encoding signal into a composite multi-track signal according to the time division multiplexing scheduling plan;
[0053] a track decomposing module, configured to decode the composite multi-track signal according to the target encoding mode and the time division multiplexing scheduling plan to obtain a recognizable target multi-language audio signal.
[0054] In a third aspect, the present application provides an electronic device for multi-language recognition, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the AI-based multi-language concurrent recognition method when executing the computer program.
[0055] In a fourth aspect, the present application provides a computer readable storage medium, which includes a computer program, and the computer program implements the steps of the AI-based multi-language concurrent recognition method when executed by a processor.
[0056] The application obtains a multi-language audio signal, trains the multi-language audio signal according to an AI neural network model to obtain a language recognition information set, determines a target coding mode according to the language recognition information set, encodes the language recognition information set according to the target coding mode to obtain a target coding signal, determines a time division multiplexing scheduling plan according to the target coding signal, schedules the target coding signal into a composite multi-audio track signal according to the time division multiplexing scheduling plan, transmits the composite multi-audio track signal in a single channel, decodes the composite multi-audio track signal after receiving the composite multi-audio track signal to obtain a recognizable target multi-language audio signal, and realizes concurrent recognition and transmission of the multi-language audio signal in the single channel. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 A flowchart of a multi-language concurrent recognition method based on AI is shown for an embodiment of the application.
[0058] Figure 2 A method flowchart executed by a language recognition module in a multi-language concurrent recognition system based on AI is shown for an embodiment of the application.
[0059] Figure 3 A method flowchart executed by a signal coding module in a multi-language concurrent recognition system based on AI is shown for an embodiment of the application.
[0060] Figure 4 A method flowchart executed by a multiplexing scheduling module in a multi-language concurrent recognition system based on AI is shown for an embodiment of the application.
[0061] Figure 5 A method flowchart executed by an audio track synthesis module in a multi-language concurrent recognition system based on AI is shown for an embodiment of the application.
[0062] Figure 6 A method flowchart executed by an audio track decomposition module in a multi-language concurrent recognition system based on AI is shown for an embodiment of the application.
[0063] Figure 7 A module schematic diagram of a multi-language concurrent recognition system based on AI is shown for an embodiment of the application.
[0064] Figure 8 A structural schematic diagram of an electronic device for multi-language recognition is shown for an embodiment of the application. DETAILED DESCRIPTION
[0065] The terminology used in the following embodiments of the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the description of the application, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0066] Hereinafter, the terms "first", "second" are only for the purpose of description, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0067] Referring to Figure 1 As shown in the figure, the AI-based multi-language concurrent recognition method shown in the embodiments of the present application comprises the following steps.
[0068] S11, acquiring a multi-language audio signal.
[0069] In the present application, the multi-language audio signal can contain at least 16 different language signals.
[0070] S12, training the multi-language audio signal according to an AI neural network model to obtain a language recognition information set.
[0071] The AI neural network model is a neural network model established based on a large amount of multi-language audio data, including language samples of various languages. According to the AI neural network model, the model training of the multi-language audio signal can identify the language category of each language in the multi-language audio signal and the start and end time of the language signal.
[0072] By adopting the above technical solution, the AI neural network model pre-set can be used to train the multi-language audio signal, and the signal category, signal start and end time, etc. of each audio signal in the multi-language audio signal can be obtained, realizing the effective differentiation of the signals in the multi-language audio signal.
[0073] S13, determining a target encoding mode according to the language recognition information set.
[0074] The encoding mode includes frequency encoding, wavelength characteristic encoding, and combined encoding of frequency encoding and wavelength characteristic encoding. The frequency encoding refers to a way of encoding information by changing the frequency characteristics of a signal. For each identified language type, it is mapped to a specific frequency range, so that the audio signals of different languages have obvious distinguishability in the frequency domain, and even if they are transmitted in the same channel, they can be identified and decoded through their respective exclusive frequency bands. Wavelength characteristic encoding refers to a special time domain encoding method, which uses the start and end time information of the language signal to design specific time domain pulse shapes, periodic patterns, or symbol structures, so that different languages have differences in time domain performance, facilitating subsequent identification and decoding.
[0075] In the language recognition information set, based on the language type in each language recognition information, the encoding mode of the corresponding language recognition information can be determined, and according to the language start and end time, the encoding mode of each corresponding language recognition information can be sorted to obtain the target encoding mode.
[0076] It should be noted that each language recognition information can be frequency encoded, wavelength characteristic encoded, and combined encoded, and based on different language recognition information, there is an optimal encoding mode. For example,
[0077] If the language recognition information of language A has obvious speech frequency characteristics, it is preferentially set to frequency encoding; if the language recognition information of language A has specific optical wavelength characteristics, it is preferentially set to wavelength characteristic encoding. If the language recognition information of language A has both speech frequency characteristics and optical wavelength characteristics, it can be set to combined encoding.
[0078] By adopting the above technical solution, the encoding mode of each language recognition information in the language recognition information set can be determined, and the target encoding mode of the entire set can be determined according to the encoding mode of the language recognition information. Based on the target encoding mode, multiple language recognition information can be effectively processed, and the language recognition information set can be encoded into a format that can be recognized and parsed according to the target encoding mode.
[0079] In an optional embodiment, the determining the target encoding mode according to the language recognition information set comprises:
[0080] Obtaining the language type and language start and end time of each language recognition information in the language recognition information set;
[0081] According to the language type, data matching is performed in the encoding strategy database to determine the encoding mode of each language recognition information, and the encoding mode includes frequency encoding, wavelength characteristic encoding, and combined encoding;
[0082] According to the language start and end time, the time sequence position of the encoding mode is determined.
[0083] According to the time sequence position, the encoding mode is sorted to obtain a target encoding mode.
[0084] The encoding strategy database is a database defined in advance in the system according to actual requirements such as language characteristics, system performance requirements, spectrum resources, application scenarios, etc. The encoding strategy database records the priority encoding mode corresponding to each language, and the backup encoding scheme under specific conditions such as spectrum resource shortage and similar language characteristics.
[0085] When the system processor receives the language recognition information set, the language category and the language start and end time of each language recognition information are obtained, the encoding mode of each language recognition information is obtained by querying the encoding strategy database according to the language category, and the encoding mode is time-sequentially arranged based on the corresponding language start and end time to determine the corresponding time sequence position, thereby obtaining the target encoding mode.
[0086] By adopting the above technical solution, the system processor can determine the corresponding encoding mode according to each language recognition information in the language recognition information set, ensure that each language recognition information is matched to the best encoding mode, and improve the encoding efficiency.
[0087] S14, encoding processing is performed on the language recognition information set according to the target encoding mode to obtain a target encoding signal.
[0088] According to the target encoding mode, the language recognition information set is encoded and processed according to the time sequence position described above, and the language recognition information is mapped to a specific frequency encoding or wavelength characteristic to obtain a target encoding signal that can be transmitted in a single channel.
[0089] By adopting the above technical solution, the language recognition information is encoded and processed by using the frequency encoding or wavelength characteristic encoding or the combination of the two, to obtain a target encoding signal that can be transmitted without interference in a single communication channel. The multi-language audio signal is converted into a target encoding signal, so that the corresponding audio signal can be recognized by the system, and the data transmission burden of the single channel is reduced, and the transmission efficiency is improved.
[0090] In an optional embodiment, the encoding processing on the language recognition information set according to the target encoding mode to obtain a target encoding signal includes:
[0091] A first recognition information set in which the encoding mode of the language recognition information set is the frequency encoding is obtained, and the following operations are performed on the first recognition information set:
[0092] Step A1, obtaining a preset frequency band, performing frequency band allocation on the first set of identification information based on the preset frequency band, and determining a target frequency band set;
[0093] Step A2, mapping the first set of identification information to the target frequency band set to obtain a first set of encoded signals;
[0094] Obtaining a second set of identification information in the language identification information set, the encoding mode of which is the wavelength characteristic encoding, and performing the following operations on the second set of identification information:
[0095] Step B1, obtaining a preset wavelength band, performing wavelength allocation on the second set of identification information based on the preset wavelength band, and determining a target wavelength band set;
[0096] Step B2, mapping the second set of identification information to the target wavelength band set to obtain a second set of encoded information;
[0097] Obtaining a third set of identification information in the language identification information set, the encoding mode of which is the set encoding, and performing the following operations on the third set of identification information:
[0098] Step C1, obtaining a frequency encoding region and a wavelength characteristic encoding region in the third set of identification information;
[0099] Step C2, performing the frequency encoding on the frequency encoding region and performing the wavelength characteristic encoding on the wavelength characteristic encoding region to obtain a third set of encoded information;
[0100] According to the first set of encoded signals, the second set of encoded signals and the third set of encoded signals, the target encoded signal is obtained.
[0101] Wherein, the first set of identification information refers to the language identification information in the language identification information set, the encoding mode of which is frequency encoding, based on frequency encoding, the specific steps of step A1 and step A2 are as follows:
[0102] Step A1 is performed, a preset frequency band is obtained, and a series of non-overlapping frequency bands are pre-set according to the number of supported languages by the system. According to the first set of identification information, the entire available frequency band is divided into a plurality of continuous or discontinuous sub-bands, each sub-band corresponds to a language, and the target frequency band set is composed.
[0103] In step A2, the first language identification information is mapped to a target frequency band set. According to the language category and start and end time in the first language identification information, the signal coding module converts the first language identification information to its corresponding frequency band according to a preset frequency band allocation scheme. For example, English audio is converted to a 3 kHz-5 kHz frequency band, French audio is converted to a 6 kHz-8 kHz frequency band, and so on.
[0104] In addition, in the processing process, the time domain language signal is also converted to the frequency domain by using a digital signal processing technology (fast Fourier transform, FFT), gain adjustment or filtering operation is performed on the specified frequency band, and the energy of the language signal is ensured to be concentrated in the corresponding frequency band, and the remaining frequency band is attenuated to an acceptable noise level. The encoded language frequency band signals are subjected to inverse Fourier transform (IFFT) to restore them to the time domain, and then superimposed to form a composite multi-language signal. The signal looks mixed together in the time domain, but the language signals are clearly separated in the frequency domain.
[0105] Through steps A1 and A2, the first identification information set is converted to a first coded signal set by frequency coding.
[0106] The second identification information set refers to the language identification information in the language identification information set that is encoded by wavelength characteristic coding. Similarly, the corresponding second identification information is mapped to a specific wavelength region to obtain a second coded signal set. The encoding process is not described in detail here.
[0107] It should be noted that the wavelength characteristic coding involves coding mode design when mapped to the preset wavelength band, that is, a unique wavelength band is defined for each language, such as a specific pulse width, pulse interval, symbol sequence, etc. When the language category and start and end time in the second identification information set are identified, the signal coding module performs time domain transformation on the original language signal according to the target wavelength band set corresponding to the second identification information set. For example, English audio signals can be converted into a pulse sequence with a specific width and interval, and French audio signals can be converted into another pulse sequence.
[0108] The third identification information set refers to the language identification information in the language identification information set that is encoded by both frequency coding and wavelength characteristic coding.
[0109] The frequency coding region and the wavelength characteristic coding region are obtained through step C1. The determination of the two regions is based on specific language characteristics and frequency resource conditions. Based on language characteristics, it is assumed that in 16 languages, the characteristics of language A and language B are similar. Only relying on frequency coding is not enough to effectively distinguish between the two. In this case, language A and language B (or even more languages with similar characteristics) can be distinguished by using a coding method that combines frequency coding and wavelength characteristics. That is, in addition to assigning different frequency bands to them in the frequency domain, specific wavelength characteristics are added in the time domain to enhance their distinguishability in a single channel. Based on frequency resources, here it refers to the frequency resources of the communication channel, not the frequency resources of the server. For example, when half of the languages have occupied most of the available frequency spectrum due to their characteristics, frequency requirements, etc., the remaining half of the languages may not be able to be effectively distinguished if they continue to use frequency coding. The half of the languages can be coded using wavelength characteristic coding, using the differences in the time domain to distinguish them without occupying additional frequency spectrum resources.
[0110] The frequency coding region and the wavelength coding region in the third identification information are determined based on language characteristics and frequency resources. The corresponding coding regions are subjected to corresponding coding steps to obtain a third set of coded signals.
[0111] According to the obtained first set of coded signals, the second set of coded signals, and the third set of coded signals, each coded signal is embedded with synchronization information (synchronization header, clock signal) according to its start and end time, so that the coded language signals are accurately superimposed in order of start and end time to obtain a target coded signal, wherein the language coded signals are interleaved in a specific time domain coding mode.
[0112] By using the above technical solutions, the corresponding coding method is determined based on different language identification information, and the corresponding coding mapping is performed on the language identification information set to obtain the target coded signal. The multi-language audio signal is transmitted in the form of the target coded signal on a single channel, and the corresponding language information of the language signal can be accurately identified, ensuring efficient and accurate identification and transmission of multiple languages on a single channel.
[0113] S15, determining a time division multiplexing scheduling plan according to the target coded signal.
[0114] The time-multiplexing scheduling plan refers to scheduling the target coded signals in time sequence on the same channel to ensure that the signal transmission does not conflict. In the time-multiplexing scheduling plan, the time of the communication channel is divided into intervals, and each interval is referred to as a time slice. Each language signal is allocated one or more time slices to determine the transmission time of the language signal on the communication channel.
[0115] By adopting the above technical solution, the time slices are dynamically allocated on a single communication link, so that different language signals can be alternately transmitted according to a preset time sequence, thereby avoiding mutual interference between signals and ensuring that each language can be correctly transmitted and decoded, thereby improving communication efficiency.
[0116] In an optional embodiment, the time-multiplexing scheduling plan is determined according to the target coded signals, and includes:
[0117] The language identification information time slice of each language identification information is determined according to the language start and end time.
[0118] The language identification information priority of each language identification information is determined according to the language type.
[0119] The time-multiplexing scheduling plan is determined according to the language identification information time slice and the language identification information priority.
[0120] According to the start and end time of each language identification information, the system processor can determine the time slice length of the corresponding language identification information. According to the language start and end time of each language identification information, the entire communication time can be divided into time slices. Each time slice corresponds to a fixed time period and is used to transmit the corresponding language identification information, thereby ensuring that each language identification information has sufficient time to be transmitted on the communication channel. According to the language type, the priority of different language types is preset in the processor of the system. For example, the transmission priority of English or Russian is high, and the priority of the corresponding language identification information is high.
[0121] By adopting the above technical solution, the language identification information time slice and the priority are combined to formulate the time-multiplexing scheduling plan, and the language identification information is transmitted in sequence according to the time slice, and the transmission order is adjusted according to the priority, so as to ensure that important information is transmitted in priority.
[0122] S16, scheduling the target coded signals into a composite multi-track signal according to the time-multiplexing scheduling plan.
[0123] When the multi-track synthesis module of the system processor obtains the target coded signals, the target coded signals are combined into a composite multi-track signal based on the determined time-multiplexing scheduling plan, and the composite multi-track signal is transmitted on a single communication channel by executing the scheduling instruction by the system scheduler.
[0124] By adopting the technical scheme, the target coded signals are combined into a composite multi-track signal based on the time division multiplexing scheduling plan, and orderly transmission of multiple signals on a single communication channel is realized, so as to avoid confusion between signals.
[0125] In an optional embodiment, the scheduling of the target coded signals into a composite multi-track signal according to the time division multiplexing scheduling plan comprises:
[0126] determining a time slice window of the target coded signal according to the language recognition information time slice;
[0127] determining a priority of the time slice window according to the language recognition information priority;
[0128] scheduling according to the priority of the time slice window to obtain the composite multi-track signal.
[0129] The system processor divides the transmission time of the target coded signal into corresponding time slice windows according to the language recognition information time slice. Each time slice window corresponds to a language recognition information time slice, and a priority of each time slice window is determined according to the language recognition information priority.
[0130] After the priority of each time slice window is determined, the time slice windows are scheduled according to the priority, and the time slice window with a high priority is transmitted before the time slice window with a low priority. The target coded signals transmitted in the scheduling order are combined to form a composite multi-track signal.
[0131] By adopting the technical scheme, the coded signals of different language recognition information can be reasonably combined in the composite multi-track signal according to the priority order, to ensure the orderliness of transmission and the execution of the priority, effectively manage and schedule the transmission of language recognition information, ensure that the language information with a high priority is processed and transmitted in time, and improve the efficiency and reliability of the communication system.
[0132] S17, decoding processing the composite multi-track signal according to the target coding mode and the time division multiplexing scheduling plan to obtain an identifiable target multi-language audio signal.
[0133] The decoding processing is performed in a multi-track decoding module, that is, the composite multi-track signal is transmitted in a single channel, and after the transmission is completed, the composite multi-track signal is received by the multi-track decoding module. According to the synthesis route of the composite multi-track signal, corresponding decoding processing is performed, and an initial identifiable target multi-language audio signal can be obtained.
[0134] It should be noted that the specific decoding manner is mainly based on the mapping relationship for decoding operation. In the multi-language concurrent recognition and transmission system, the synthesis part (MTM-Syn) combines the coded language signals into a composite multi-audio track signal according to the instruction of the time division multiplexing scheduler. A mapping relationship is established in this process, that is, the coded information of each language is associated with a specific time slice, frequency range or wavelength characteristic. The decomposition part (MTM-Dec) receives the composite signal at the receiving end and decodes according to the pre-set mapping relationship to restore the original target multi-language audio stream. The target multi-language audio signal can be output to the corresponding language audio processing unit or the playback unit for processing.
[0135] By adopting the above technical solution, the multi-language audio signal is converted into a composite multi-audio track signal for transmission in a single channel, realizing the transmission of multiple language signals in a single channel, and providing decoding of the composite multi-audio track signal to obtain the target multi-language audio signal, realizing effective recognition of the multi-language audio signal.
[0136] In an optional embodiment, the decoding processing of the composite multi-audio track signal according to the target coding manner and the time division multiplexing scheduling plan to obtain the recognizable target multi-language audio signal includes:
[0137] Timing decoding is performed according to the time division multiplexing scheduling plan to obtain first decoding information.
[0138] The first decoding information is frequency decoded to obtain second decoding information.
[0139] The second decoding information is wavelength decoded to obtain third decoding information.
[0140] The third decoding information is executed for multi-language audio reconstruction to obtain the target multi-language audio signal.
[0141] According to the time division multiplexing scheduling plan, the composite multi-audio track signal is sliced in the order of time slice windows, and the signals in each time slice window are separated to obtain first decoding information, wherein each time slice window corresponds to an encoded language signal.
[0142] The frequency encoded signal in the first encoded signal is obtained, and the corresponding frequency encoded signal is extracted from the first decoding information according to the filter and the matching pursuit technology, and the corresponding mapped plane is restored to the original language recognition information to obtain the second decoding signal. For example, assuming that English is mapped to the frequency band of 3kHz-5kHz, the coded information of English is recovered from this part of the frequency spectrum of the composite signal by the filter and the matching pursuit technology.
[0143] Similarly, through the filter and the matching pursuit technology, the second decoding is carried out for the wavelength decoding, and the encoded signal mapped to the wavelength band is restored to the original language recognition information. It should be noted that for the signal using the frequency encoding and the wavelength characteristic encoding at the same time, through the above frequency decoding and wavelength decoding, the original language recognition signal, i.e., the third decoding information, can also be restored.
[0144] Finally, the third decoding information is recombined according to the arrangement order of the original multi-language audio signal to obtain the same multi-language audio signal as the sending end, i.e., the target multi-language audio signal.
[0145] Through the above technical solution, the entire decoding process strictly follows the mapping relationship established by the sending end, and based on the time division multiplexing scheduling, frequency encoding, wavelength decoding, and signal recombination, the processing and decoding of the composite multi-audio track signal can be effectively realized, the accurate separation and restoration of the language components in the composite signal are ensured, the closed-loop working process of the multi-language concurrent recognition and transmission system is realized, and the technical problem of concurrent recognition and transmission of multiple languages in a single channel is solved.
[0146] In an optional embodiment, the scheme further includes:
[0147] The channel state is monitored in real time, and the target encoding mode and the time division multiplexing scheduling plan are adjusted according to the channel state.
[0148] In the actual running process of the system, the system processor monitors and evaluates the current system state and channel condition in real time, such as the language characteristic difference, the available frequency range of the channel, the actual demand factors such as the channel bandwidth, and the like. When the frequency resources in the channel are tight, the system processor will dynamically adjust the pre-defined encoding mode. For example, according to the language characteristic difference and the current frequency spectrum resource condition, the wavelength characteristic encoding is used for supplement or replacement for some languages to enhance the distinguishing ability or save the frequency spectrum resources.
[0149] After the target encoding mode is modulated according to the channel state, the time division multiplexing scheduling plan is updated again based on the modulated target encoding mode, and the dynamic resource modulation in the transmission process of the signal is realized.
[0150] Through the above technical solution, the current system state and channel condition are monitored and evaluated in real time, the encoding mode is flexibly adjusted according to the actual situation, the system can efficiently utilize the frequency spectrum resources in different environments, the most suitable encoding mode and time division multiplexing scheduling plan are selected, different language characteristics and frequency spectrum resource conditions can be effectively coped with, the accuracy and reliability of the language recognition information are improved, and efficient and stable language recognition and transmission services are provided for users.
[0151] In summary, the application trains multiple language audio signals through an AI neural network model to obtain a language recognition information set, determines a target coding mode based on the language information set, encodes the language recognition information set according to the target coding mode to obtain a target coded signal, schedules the target coded signal into a composite multi-audio track signal according to a time division multiplexing scheduling plan, transmits the composite multi-audio track signal in a single channel, decodes the composite multi-audio track signal after receiving the composite multi-audio track signal to obtain a recognizable target multi-language audio signal, and solves the technical problem that it is difficult to concurrently recognize and transmit multiple languages in a single channel in the prior art.
[0152] In addition, in order to better understand the above technical solutions, the above technical solutions are described in combination with Figures 2 to 6 The above technical solutions are described in combination with
[0153] S21, a signal acquisition module;
[0154] S11 is used to acquire external multi-language audio signals and transmit the multi-language audio signals to the language recognition module S22 to S26.
[0155] S22, a multi-language audio signal;
[0156] S22 is used to transmit the multi-language audio signal.
[0157] S23, an AI neural network model;
[0158] The AI neural network model is used to train the composite multi-language in S22.
[0159] S24, a recognition result;
[0160] S24 obtains a language recognition information set, which includes a language type and language start and end times.
[0161] S25, a signal coding module (SC);
[0162] The language recognition information set is transmitted to the signal coding module to perform S31 to S35.
[0163] S26, an explanation of the recognition result of S24, which indicates that the recognition result is obtained by training the multi-language audio signal based on an AI neural network model using deep learning.
[0164] Based on the above signal recognition module, the input is a multi-language audio signal, the processing mode is real-time recognition using a deep learning model, the output is a recognition result of each language (a language type, a start time, and an end time), and the output is the recognition result sent to a signal coding module, wherein S31 to S35 are used to perform S13 and S14.
[0165] S31, SC (Signal Coding Module);
[0166] Connected to S25 above, the signal encoding module acquires the language recognition information set;
[0167] S32, frequency coding or wavelength characteristic coding;
[0168] Based on the language type and start and end time of each language recognition information in the language recognition information set, the encoding method of each language is determined, and frequency encoding, wavelength characteristic encoding, or a combination of both is performed to obtain the target encoded signal.
[0169] S33, target encoded signal;
[0170] Transmit the target encoded signal to S34:
[0171] S34, TDM (Time Division Multiplexing, multiplexing scheduling module);
[0172] S35 provides encoded language signals to be scheduled;
[0173] The supplementary explanation for S32 indicates that multiple language signals to be scheduled are obtained after encoding.
[0174] Based on the above signal encoding module, the input is the recognition result. The processing method involves frequency encoding or wavelength characteristic encoding for each language according to the recognition result. The output is the sending of the target encoded signal to the multiplexing scheduling module, where S41 to S45 are used to execute S15.
[0175] S41, TDM (Time Division Multiplexing, multiplexing scheduling module);
[0176] Connected to the above S35, the multiplexing scheduling module obtains the target code;
[0177] S42, Time-sharing multiplexing scheduling plan;
[0178] Based on the language type of each language recognition information in the language recognition information set, determine the language recognition information time slice for each language recognition information; based on the start and end times of each language recognition information in the language recognition information set, determine the language recognition information priority for each language recognition information; based on the language recognition information priority mentioned in the language recognition information time slice, determine the time-division multiplexing scheduling plan.
[0179] S43, Time Slice Allocation;
[0180] Time slices are allocated according to the reuse scheduling plan.
[0181] S44, multi-track decomposition module (MTM-Dec);
[0182] S45, providing time allocation information to synthesize the composite multi-track signal.
[0183] For step S42, based on the time allocation information, time slice allocation can be performed to obtain the composite multi-track signal through time division multiplexing plan.
[0184] Based on the above multiplexing scheduling module, input: target encoded signal. Processing method: make a time division multiplexing plan, allocate time slices for each language. Output: scheduling plan to multi-track synthesis module; scheduling instruction feedback to signal encoding module. It should be noted that K represents Figure 5 the multi-track synthesis module in S16.
[0185] S51, K (used to represent the multi-track synthesis module);
[0186] S52, composite multi-track signal;
[0187] According to the language recognition information, the time slice window of the target encoded signal is determined; according to the priority of the language recognition information, the priority of the time slice window is determined; according to the priority of the time slice window, the composite multi-track signal is obtained.
[0188] S53, communication channel;
[0189] The composite multi-track signal is transmitted in a single communication channel.
[0190] S54, multi-track decomposition module (MTM-Dec);
[0191] S55, restoring the original multi-language audio signal.
[0192] For step S52, based on the composite multi-track signal, the original multi-language audio signal can be restored.
[0193] Based on the above multiplexing scheduling module, input: time division multiplexing scheduling plan and target encoded signal. Processing method: according to the time division multiplexing scheduling plan, the target encoded language signal is scheduled to the composite multi-track signal. Output: the composite multi-track signal to the communication channel to the multi-track decomposition module. It should be noted that N represents Figure 6 the multi-track decomposition module in S17.
[0194] S61, N (multi-track decomposition module);
[0195] S62, target multi-language audio signal;
[0196] According to the time division multiplexing scheduling plan, time sequence decoding is performed to obtain first decoding information; frequency decoding is performed on the first decoding information to obtain second decoding information; wavelength decoding is performed on the second decoding information to obtain third decoding information; and multi-language audio reconstruction is performed on the third decoding information to obtain a target multi-language audio signal.
[0197] S63, end the current multi-language concurrent recognition.
[0198] In addition, referring to Figure 7 As shown in the figure, the AI-based multi-language concurrent recognition system 70 provided by the embodiment of the application includes a signal acquisition module 701, a language recognition module 702, a signal encoding module 703, a multiplexing scheduling module 704, a track synthesis module 705, a track decomposition module 706, and a real-time detection module 707.
[0199] The signal acquisition module 701 is configured to acquire a multi-language audio signal.
[0200] The language recognition module 702 is configured to train the multi-language audio signal according to an AI neural network model to obtain a language recognition information set.
[0201] The signal encoding module 703 is configured to determine a target encoding mode according to the language recognition information set.
[0202] The signal encoding module 703 is further configured to perform encoding processing on the language recognition information set according to the target encoding mode to obtain a target encoded signal.
[0203] The multiplexing scheduling module 704 is configured to determine a time division multiplexing scheduling plan according to the target encoded signal.
[0204] The track synthesis module 705 is further configured to schedule the target encoded signal into a composite multi-track signal according to the time division multiplexing scheduling plan.
[0205] The track decomposition module is configured to perform decoding processing on the composite multi-track signal according to the target encoding mode and the time division multiplexing scheduling plan to obtain a target multi-language audio signal that can be recognized.
[0206] Based on the above module, the multi-language audio signal is trained through the AI neural network model, and a language recognition information set can be obtained; a target coding mode is determined based on the language information set; the language recognition information set is coded according to the target coding mode to obtain a target coded signal; and the target coded signal is scheduled as a composite multi-audio track signal according to a time division multiplexing scheduling plan; the composite multi-audio track signal is transmitted in a single channel, and after receiving the composite multi-audio track signal, the composite multi-audio track signal is decoded to obtain a recognizable target multi-language audio signal, thereby solving the technical problem that it is difficult to concurrently recognize and transmit multiple languages in a single channel in the prior art.
[0207] Referring to Figure 8 As shown in FIG. 1, it is a structural schematic diagram of a printer according to an embodiment of the present application. In the preferred embodiment of the present application, the electronic device 8 comprises a memory 81, at least one processor 82, and at least one communication bus 83.
[0208] Those skilled in the art should understand that Figure 8 The structure of the electronic device shown in the figure is not a limitation of the embodiments of the present application, and can be a bus structure or a star structure. The printer can further comprise more or less other hardware or software, or different component arrangements.
[0209] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium is, for example, a non-volatile memory, such as a magnetic medium (e.g., a hard disk, a floppy disk, and a magnetic tape), an optical medium (e.g., a CDROM disk and a DVD), a magneto-optical medium (e.g., an optical disk), and a hardware device specially constructed for storing and executing computer executable instructions (e.g., a read-only memory (ROM), a random access memory (RAM), a flash memory, etc.). The computer readable storage medium stores computer executable instructions. The computer readable storage medium can execute the computer executable instructions by one or more processors or processing devices to implement the aforementioned AI-based multi-language concurrent recognition method.
[0210] In addition, it can be understood that the foregoing embodiments are only exemplary descriptions of the present application, and the technical solutions of the embodiments can be arbitrarily combined and used without conflict in technical features, contradiction in structure, and violation of the purpose of the present application.
[0211] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the system or apparatus embodiments described above are merely illustrative, for example, the division of units is merely a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, systems or units, which can be electrical, mechanical or other forms.
[0212] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0213] In addition, the functional units / modules in each embodiment of the present application can be integrated into one processing unit / module, or each unit / module can be physically present alone, or two or more units / modules can be integrated into one unit / module. The integrated unit / module can be realized in the form of hardware or hardware plus software functional unit / module.
[0214] The integrated unit / module realized in the form of software functional unit / module can be stored in a computer readable storage medium. The software functional unit stored in a storage medium includes a plurality of instructions for causing one or more processors of a computer device (which can be a personal computer, a server, or a network device, etc.) to execute part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0215] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An AI-based multi-language concurrent recognition method, characterized by, The method comprises the following steps: acquiring a multi-language audio signal; training the multi-language audio signal according to an AI neural network model to obtain a language recognition information set, which is composed of the language category of each language in the multi-language audio signal and the start and end time of the language signal; determining a target encoding mode according to the language recognition information set; encoding the language recognition information set according to the target encoding mode to obtain a target encoding signal; determining a time division multiplexing scheduling plan according to the target encoding signal; scheduling the target encoding signal into a composite multi-audio track signal according to the time division multiplexing scheduling plan; decoding the composite multi-audio track signal according to the target encoding mode and the time division multiplexing scheduling plan to obtain an identifiable target multi-language audio signal. 2.The AI-based multi-language concurrent recognition method of claim 1, wherein, The step of determining the target encoding mode according to the language recognition information set comprises the following steps: acquiring the language type and the start and end time of each language recognition information in the language recognition information set; determining the encoding mode of each language recognition information in the encoding strategy database according to the language type, wherein the encoding mode comprises frequency encoding, wavelength characteristic encoding and combined encoding; determining the time sequence position of the encoding mode according to the start and end time of the language; sorting the encoding mode according to the time sequence position to obtain the target encoding mode. 3.The AI-based multi-language concurrent recognition method of claim 2, wherein, The step of encoding the language recognition information set according to the target encoding mode to obtain the target encoding signal comprises the following steps: acquiring a first recognition information set in the language recognition information set, wherein the encoding mode of the first recognition information set is the frequency encoding, and performing the following operations on the first recognition information set: Step A1: acquiring a preset frequency band, performing frequency band allocation on the first recognition information set based on the preset frequency band, and determining a target frequency band set; Step A2: mapping the first recognition information set to the target frequency band set to obtain a first encoding signal set; acquiring a second recognition information set in the language recognition information set, wherein the encoding mode of the second recognition information set is the wavelength characteristic encoding, and performing the following operations on the second recognition information set: Step B1: acquiring a preset wavelength band, performing wavelength allocation on the second recognition information set to determine a target wavelength set; Step B1: acquiring a preset wavelength band, performing wavelength allocation on the second recognition information set based on the preset wavelength band to determine a target wavelength band set; Step B2: mapping the second recognition information set to the target wavelength band set to obtain a second encoding information set; acquiring a third recognition information set in the language recognition information set, wherein the encoding mode of the third recognition information set is the combined encoding, and performing the following operations on the third recognition information set: Step C1: acquiring the frequency encoding area and the wavelength characteristic encoding area in the third recognition information set; Step C2: performing the frequency encoding on the frequency encoding area and performing the wavelength characteristic encoding on the wavelength characteristic encoding area to obtain a third encoding information set; combining the first encoding signal set, the second encoding signal set and the third encoding signal set to obtain the target encoding signal.
4. The AI-based multi-language concurrent recognition method of claim 3, wherein, The determining the time division multiplexing scheduling plan according to the target coded signal comprises: determining a language identification information time slice of each language identification information according to the language start and end time; determining a language identification information priority of each language identification information according to the language type; determining the time division multiplexing scheduling plan according to the language identification information time slice and the language identification information priority.
5. The AI-based multi-language concurrent recognition method according to claim 4, characterized in that, The scheduling the target coded signal into a composite multi-audio track signal according to the time division multiplexing scheduling plan comprises: determining a time slice window of the target coded signal according to the language identification information time slice; determining a time slice window priority according to the language identification information priority; scheduling according to the time slice window priority to obtain the composite multi-audio track signal.
6. The AI-based multi-language concurrent recognition method of claim 1, wherein, The decoding processing of the composite multi-audio track signal according to the target coding mode and the time division multiplexing scheduling plan to obtain the recognizable target multi-language audio signal comprises: time sequence decoding according to the time division multiplexing scheduling plan to obtain first decoding information; frequency decoding of the first decoding information to obtain second decoding information; wavelength decoding of the second decoding information to obtain third decoding information; multi-language audio reconstruction of the third decoding information to obtain the target multi-language audio signal.
7. The AI-based multi-language concurrent recognition method according to any one of claims 1 to 5, characterized in that, The method further comprises: real-time monitoring of a channel state, and adjusting the target coding mode and the time division multiplexing scheduling plan according to the channel state.
8. An AI-based multi-language concurrent system, characterized by, The method for concurrent multi-language recognition based on AI according to any one of claims 1-7 comprises: a signal acquisition module for acquiring a multi-language audio signal; a language recognition module for training the multi-language audio signal according to an AI neural network model to obtain a language identification information set, the language identification information set being composed of a language type of each language in the multi-language audio signal and a start and end time of the language signal; a signal coding module for determining a target coding mode according to the language identification information set; the signal coding module is further configured to encode process the language identification information set according to the target coding mode to obtain a target coded signal; a multiplexing scheduling module for determining a time division multiplexing scheduling plan according to the target coded signal; an audio track synthesis module for scheduling the target coded signal into a composite multi-audio track signal according to the time division multiplexing scheduling plan; an audio track decomposition module for decoding processing of the composite multi-audio track signal according to the target coding mode and the time division multiplexing scheduling plan to obtain a recognizable target multi-language audio signal. 9.A multi-language recognition electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the multi-language recognition method based on AI according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, An electronic device comprising a memory having stored thereon executable code that, when executed by a processor of the electronic device, causes the processor to perform steps of a method of AI-based concurrent multilingual recognition as claimed in any of claims 1 to 7.
Citation Information
Patent Citations
Multiple access communications system and method using code and time division
CA2246535A1
Meeting implementation method, device, equipment and system and computer readable storage medium
CN108076306A