A communication information sorting system based on the analysis of the historical parameters of large models

The communication information sorting system enhances walkie-talkie clarity and efficiency by analyzing audio spectra and sorting audio frames based on user habits, addressing noise-induced communication challenges and improving user experience.

CN119210503BActive Publication Date: 2025-07-15SHENZHEN HONGKETE ELECTRONIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411307511.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-07-15
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

When the intercom is exploring in the field, the noise of the communication environment affects the communication voice quality, resulting in increased communication difficulty and increased transmission tasks, and the user location cannot be determined, which poses a safety hazard.

Method used

The communication information sorting system based on the analysis of the historical parameter of the big model is adopted. User audio is collected through multiple sets of microphones, the audio source direction is identified, and the audio spectrum marking and segmentation processing is performed. The audio frame sorting queue is generated based on information entropy recognition, and audio frame sorting and reorganization are performed to eliminate noise and improve communication clarity.

Benefits of technology

It realizes the clarity and simplicity of the audio transmission of the walkie-talkie, reduces noise, improves the user's communication experience, and ensures effective communication of the walkie-talkie in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119210503B_ABST
    Figure CN119210503B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and specifically relates to a communication information sorting system based on the analysis of historical parameters of a large model, including a control terminal, a receiving layer, a sorting layer, and an interaction layer; the control terminal, which is the main control end of the system, is used to decide the start and stop of the system and the start and stop of each module in the system; the audio appearing at the walkie-talkie audio sending end is received by the receiving layer, and is used to obtain the corresponding audio spectrum based on the received audio, identify the audio source direction through the audio spectrum, and mark the audio and the audio spectrum based on the identified audio source direction. The sorting layer receives the audio and the audio spectrum. The present invention performs operations such as splitting the corresponding audio spectrum of the audio and converting text information, so as to achieve the goal of finally reducing the audio frames that do not contain information in the audio, and finally making the audio transmitted by the walkie-talkie clearer, more concise, and with less noise. Furthermore, when the user uses the walkie-talkie, this system brings a better communication experience to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a communication information sorting system based on the analysis of historical parameters of large models. Background Art

[0002] Walkie-talkie communication is a convenient two-way wireless communication method. It transmits voice signals through specific frequency bands, is easy to operate, and does not require complex settings. It is suitable for short-distance communication, such as construction sites, security patrols, outdoor activities and other scenarios. Users can press the button to speak and release it to listen, achieving instant communication.

[0003] The invention patent with the application number 202311025009.3 discloses a communication method for a walkie-talkie. The communication method is characterized in that it is applied to a communication system of a walkie-talkie. The communication system includes a starting walkie-talkie, a relay walkie-talkie and a target terminal. The number of relay walkie-talkies is N, where N is a positive integer. The communication method includes: the starting walkie-talkie obtains the device information of the starting walkie-talkie and sends the device information to the relay walkie-talkies within the first communication range. The device information at least includes location information. The first communication range is the communication range of the starting walkie-talkie. The relay walkie-talkie receives the device information and forwards the device information to the target terminal. There is the target terminal within the second communication range of at least one relay walkie-talkie. The second communication range is the communication range of the relay walkie-talkie. The target terminal receives the device information.

[0004] This application aims to solve the problem that "in some specific situations, such as during field exploration, when the physical distance between the staff holding the walkie-talkie and the control center exceeds the communication distance of the walkie-talkie, users cannot communicate with the control center through the walkie-talkie to report their location. At this time, the location of the user cannot be determined, posing a safety hazard".

[0005] However, during the voice communication process of the walkie-talkie, background noise in the communication environment will affect the quality of the communication voice, and due to the background noise in the communication environment, the communication voice data packet is also larger, so that the task volume of transmitting the communication voice by the walkie-talkie increases, and thus the communication difficulty also increases accordingly.

[0006] Therefore, we propose a communication information sorting system based on the analysis of historical parameters of large models. Summary of the Invention

[0007] In view of the above-mentioned drawbacks of the prior art, the present invention provides a communication information sorting system based on the analysis of historical parameters of large models, which solves the technical problems raised in the above background art.

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0009] A communication information sorting system based on the analysis of historical parameters of large models, including a control terminal, a receiving layer, a sorting layer and an interaction layer;

[0010] The control terminal is the main control end of the system, used to decide the start and stop of the system and the start and stop of each module in the system;

[0011] The audio that appears at the walkie-talkie audio sending end is received by the receiving layer. At the same time, the corresponding audio spectrum is obtained based on the received audio, the audio source direction is identified through the audio spectrum, and the audio and audio spectrum are marked based on the identified audio source direction. The sorting layer receives the audio and audio spectrum, and uses the audio spectrum to perform segmentation processing on the audio to obtain several groups of audio frames corresponding to different unit times. Further, the sorted remaining audio frames are fed back to the interaction layer. The interaction layer recombines the received audio frames and sends the combined audio to the walkie-talkie audio receiving end;

[0012] The sorting layer includes a phonetic translation module, a segmentation module and a sorting module. The phonetic translation module is used to obtain the received audio and the corresponding audio spectrum in the receiving layer, further select the audio and audio spectrum with marks, convert the selected audio into text information, and forward the obtained audio spectrum to the segmentation module. The segmentation module is used to receive the audio spectrum and perform segmentation processing on the audio spectrum. The sorting module is used to receive the audio frames segmented in the segmentation module and the text information conversion result of the audio, and sort the audio frames in combination with the audio frames and the text information conversion result;

[0013] When the sorting module receives the audio frames and the text information conversion result of the audio, it synchronously performs information entropy recognition on each audio frame. Based on the information entropy recognition result of the audio frame, an audio frame sorting queue is generated, and then the audio frame sorting operation is executed;

[0014] The recognition logic of the information entropy of the audio frame is expressed as:

[0015]

[0016] In the formula: H is the information entropy of the audio frame; K is the total amount of feature values in the audio frame; p j is the probability of the jth group of feature values appearing; ε j is the weight of the jth group of feature values; θ is the correction;

[0017] Among them, when identifying the information entropy of an audio frame, with the peak frequency point in the audio frame as the boundary, the length of the shortest line segment among the line segments on both sides of the peak frequency point in the audio frame is used as the truncation reference for the other line segment. Starting from the peak frequency point, a line segment equal in length to the shortest line segment is intercepted on the other line segment. Further, a segmentation accuracy is set, and based on the segmentation accuracy, the shortest line segment and the intercepted line segment among the line segments on both sides of the peak frequency point in the audio frame are segmented to obtain several groups of sub-line segments. The maximum value among the amplitude values represented by both ends of the several groups of sub-line segments is denoted as the eigenvalue in the audio frame, and the weight ε of the j-th group of eigenvalues j obeys that the closer the sub-line segment from which the eigenvalue is derived is to the peak frequency point in the audio frame, the weight ε j has a larger value, and the weight ε j > 0, and

[0018] Furthermore, the receiving layer includes a receiving module, an identifying module, and a marking module. The receiving module is used to receive the audio that appears at the audio transmitting end of the walkie-talkie and synchronously convert the received audio into an audio spectrum. The identifying module is used to obtain the audio spectrum obtained by converting the audio in the receiving module and identify the audio source direction based on the audio spectrum. The marking module is used to receive the audio source direction identification result in the identifying module and mark the audio belonging to the audio source direction and its corresponding audio spectrum to distinguish it from other audio received by the receiving module and the corresponding audio spectrum;

[0019] Among them, the receiving module is integrated by three groups of microphones. The three groups of microphones are distributed in a triangular shape, and the distance between each adjacent two groups of microphones is equal. The receiving module is installed on the walkie-talkie to receive the audio emitted by the user at the audio transmitting end of the walkie-talkie in real time and import the audio spectrum into the audio spectrum analyzer, and output the audio spectrum corresponding to the audio through the audio spectrum analyzer.

[0020] Furthermore, the audio spectrum analyzer is deployed inside the walkie-talkie. After the microphone receives the audio emitted by the user at the audio transmitting end of the walkie-talkie, it is transmitted to the audio spectrum analyzer in real time. Each group of microphones in the receiving module receives a group of audio, and each group of audio outputs its corresponding audio spectrum through the audio spectrum analyzer;

[0021] The audio source direction identification logic in the identifying module is expressed as:

[0022]

[0023] In the formula: SIMM(a,b) is the similarity between the audio spectrum a and the audio spectrum b; f p1 is the value of the frequency point with the maximum energy in the audio spectrum a; f p2 is the value of the frequency point with the maximum energy in the audio spectrum b; f maxis the value of the frequency point with the largest energy in two groups of audio spectra; ω1 is the weight; N is the total number of frequency bands obtained by dividing the audio spectrum; E i-a is the energy ratio of the i-th frequency band of audio spectrum a; E i-b is the energy ratio of the i-th frequency band of audio spectrum b;

[0024] Among them, the similarity between three groups of audio spectra is obtained based on the above formula, and the f in the calculation result of the similarity with the lowest similarity is selected from max the source audio spectrum as the reference for audio source direction recognition. The microphone of the audio source direction recognition reference source is the audio source direction.

[0025] Furthermore, the weight ω1 ≤ 0.5. Before calculating the similarity between audio spectrum a and audio spectrum b, the two groups of audio spectra are divided so that the time of each sub-audio spectrum obtained by division is one second;

[0026] Before calculating the similarity between the audio spectrum a and the audio spectrum b, the two groups of audio spectra are normalized so that the total energy of the two groups of audio spectra is 1;

[0027] The marking content of the audio spectrum and audio corresponding to the audio source direction recognition result in the marking module is a text mark, and the text mark content is: origin, and the other audio and corresponding audio spectra, that is, the remaining two groups of audio and audio spectra.

[0028] Furthermore, after receiving the audio spectrum, the segmentation module picks up all the valley-peak frequency points in the audio spectrum, and takes the spectral segments where every three adjacent frequency points are located as a group of audio frames. Among the three adjacent frequency points, there are two groups of valley frequency points and one group of peak frequency points, and the peak frequency point is between the two groups of valley frequency points. The above is recorded as the audio frame determination logic. The segmentation module segments the audio spectrum based on the audio frame determination logic to obtain several groups of audio frames, and the unit time of several groups of the audio frames is not equal;

[0029] After the segmentation module segments the audio spectrum into several groups of audio frames, the segmentation module synchronously obtains the audio obtained in the transliteration module, and based on the ratio of the respective unit times of several groups of audio frames, the audio is segmented to obtain several groups of sub-audios. After several groups of sub-audios are configured behind the corresponding audio frames, they are sent to the transliteration module to perform text information conversion operations on each group of sub-audios. Further, the converted text information is bound to the audio frames corresponding to the sub-audios where the text information source is located, and then returned to the segmentation module, and the bound text information and audio frames are sent to the sorting module based on the segmentation module.

[0030] Further, 2 > θ ≥ 1, and it follows that: the greater the length difference between the line segments on both sides of the peak frequency point in the audio frame, the greater the corrected value of θ; conversely, the smaller the corrected value of θ.

[0031] When the sorting module receives the audio frame and the conversion result of the text information of the audio, that is, the text information and the audio frame that are mutually bound, when the sorting module runs to generate the audio frame sorting queue, it sorts each audio frame in descending order according to the magnitude of the information entropy corresponding to the audio frame to generate the audio frame sorting queue, and gives priority to the audio frames with large information entropy in the audio frame sorting queue as the sorting targets.

[0032] The logic of the audio frame sorting operation in the sorting module is expressed as:

[0033]

[0034] In the formula: Q t is the determination value of whether the audio frame t is a sorting target; L t is the text information converted from the audio frame based on the transliteration module; T0 is the text information converted from the audio based on the transliteration module.

[0035] Among them, based on the above formula, the determination value of whether each audio frame in the audio frame sorting queue is a sorting target is obtained.

[0036] If the determination value corresponding to whether the audio frame is a sorting target is 1, then the audio frame is the sorting target audio frame of the sorting module, and the sorting target audio frame is deleted by the sorting module.

[0037] Further, the interaction layer includes a picking module, a recombination module, and a refreshing module. The picking module is used to pick the audio frames remaining after sorting by the sorting module in the sorting layer, further obtain the corresponding time domain of each remaining sorted audio frame in the audio spectrum, and intercept the audio with the same time domain in the audio to be recorded as sub-audio. The recombination module is used to receive all the sub-audios intercepted by the picking module, sort each group of sub-audios based on their corresponding time domains to form a new audio. The refreshing module is used to receive the new audio obtained by the operation of the recombination module, use the new audio as the audio to be transmitted by the audio sending end of the walkie-talkie, transmit it to the audio receiving end of the walkie-talkie, and refresh the system operation.

[0038] Further, when the recombination module performs sorting and recombination operations on each group of sub-audios, fade-in and fade-out processing is performed on each adjacent two groups of sub-audios.

[0039] Among them, when fade-in and fade-out processing is performed on two adjacent groups of sub-audios, the fade-in and fade-out duration follows the rule that the longer the duration of the sub-audio with the shortest time domain in two adjacent groups of sub-audios, the longer the fade-in and fade-out duration; conversely, the shorter the fade-in and fade-out duration. When fade-in and fade-out processing is performed on two adjacent groups of sub-audios, the fade-in and fade-out audio intensity is within the audio intensity threshold formed by the maximum audio intensity and the minimum audio intensity in two adjacent groups of sub-audios.

[0040] Furthermore, during the operation stage of the refresh module, the recognition result of the audio source direction recognized by the recognition module in the receiving layer during the current operation of the system is synchronously recorded. When the audio source directions recorded twice in a row are the same, when the system operates next under the control of the refresh module, the recognition module does not operate, and the previous audio source direction recognition result is applied to the marking module;

[0041] Among them, during each operation process of the system, the operation process of the system in which the recognition module does not operate is not used as the determination target for the audio source directions recorded twice in a row to be the same.

[0042] Furthermore, the transliteration module is interconnected with a segmentation module and a sorting module through a local area network. The transliteration module is interconnected with a marking module through a local area network. The marking module is interconnected with a recognition module and a receiving module through a local area network. The sorting module is interconnected with a picking module through a local area network. The picking module is interconnected with a recombination module and a refresh module through a local area network.

[0043] Adopting the technical solution provided by the present invention, compared with the known public technology, it has the following beneficial effects:

[0044] The present invention provides a communication information sorting system based on the analysis of historical parameters of a large model. During the operation of the system, by setting multiple groups of microphones, the historical usage habits of walkie-talkie users are collected. Thus, based on the historical usage habits of walkie-talkie users, the user audio received by multiple groups of microphones is selected to ensure that the system can use the clearest group of audio that conforms to the original meaning of the walkie-talkie user audio as the further operation target of the system. Then, operations such as splitting the corresponding audio spectrum of the audio and converting text information are performed to achieve the goal of finally eliminating the audio frames that do not contain information in the audio. Finally, the audio transmitted by the walkie-talkie is made clearer, more concise, and contains less noise. Furthermore, when the user uses the walkie-talkie, a better communication experience is brought to the user based on this system. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic structural diagram of a communication information sorting system based on the analysis of historical parameters of a large model;

[0047] Figure 2 It is an example diagram of the deployment pose of the receiving module on the walkie-talkie in the present invention;

[0048] Figure 3 It is an example diagram of the audio spectrum in the present invention;

[0049] Figure 4 It is a schematic diagram of the source of the eigenvalues used for audio frame segmentation, recombination logic, and audio frame information entropy calculation in the audio spectrum of the present invention. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0051] The following further describes the present invention with reference to the embodiments.

[0052] Embodiment 1:

[0053] A communication information sorting system based on the analysis of historical parameters of a large model in this embodiment, as Figure 1 shown, includes a control terminal, a receiving layer, a sorting layer, and an interaction layer;

[0054] The control terminal is the main control end of the system and is used to decide the start and stop of the system and the start and stop of each module in the system;

[0055] The audio that appears at the audio sending end of the walkie-talkie is received through the receiving layer. Meanwhile, the corresponding audio spectrum is obtained based on the received audio, and the audio source direction is identified through the audio spectrum. Based on the identified audio source direction, the audio and its corresponding audio spectrum are marked. The sorting layer receives the audio and the audio spectrum, and uses the audio spectrum to perform segmentation processing on the audio, obtaining several groups of audio frames corresponding to different unit times. Further, the sorted remaining audio frames are fed back to the interaction layer. The interaction layer reorganizes the received audio frames and sends the combined audio to the audio receiving end of the walkie-talkie;

[0056] The receiving layer includes a receiving module, an identifying module, and a marking module. The receiving module is used to receive the audio that appears at the audio sending end of the walkie-talkie, and simultaneously convert the received audio into an audio spectrum. The identifying module is used to obtain the audio spectrum obtained by converting the audio in the receiving module, and identify the audio source direction based on the audio spectrum. The marking module is used to receive the audio source direction identification result in the identifying module, and mark the audio belonging to the audio source direction and its corresponding audio spectrum to distinguish it from other audio and corresponding audio spectra received by the receiving module;

[0057] Among them, the receiving module is integrated by three groups of microphones. The three groups of microphones are distributed in a triangular shape, and the distance between each adjacent two groups of microphones is equal. The receiving module is installed on the walkie-talkie, receives the audio emitted by the user at the audio sending end of the walkie-talkie in real time, and imports the audio spectrum into the audio spectrum analyzer, and outputs the audio spectrum corresponding to the audio through the audio spectrum analyzer;

[0058] The sorting layer includes a transliteration module, a segmentation module, and a sorting module. The transliteration module is used to obtain the audio and its corresponding audio spectrum received in the receiving layer, further select the audio and audio spectrum with marks, convert the selected audio into text information, and forward the obtained audio spectrum to the segmentation module. The segmentation module is used to receive the audio spectrum and perform segmentation processing on the audio spectrum. The sorting module is used to receive the audio frames obtained by segmentation in the segmentation module and the text information conversion result of the audio, and sort the audio frames in combination with the audio frames and the text information conversion result;

[0059] When the sorting module receives the audio frames and the text information conversion result of the audio, it synchronously performs information entropy identification on each audio frame. Based on the information entropy identification result of the audio frame, an audio frame sorting queue is generated, and then the audio frame sorting operation is performed;

[0060] The identification logic of the information entropy of the audio frame is expressed as:

[0061]

[0062] In the formula: H is the information entropy of the audio frame; K is the total amount of eigenvalue in the audio frame; p j is the probability of the jth group of eigenvalues appearing; ε jis the weight of the j-th group of eigenvalues; θ is the correction;

[0063] 2 > θ ≥ 1, and it follows that: the greater the difference in the lengths of the line segments on both sides of the peak frequency point in the audio frame, the greater the value of the correction θ, and vice versa, the smaller the value of the correction θ;

[0064] When the sorting module receives the audio frame and the conversion result of the text information of the audio, that is, the text information and the audio frame that are mutually bound, when the sorting module runs to generate the audio frame sorting queue, the audio frames are sorted in descending order according to the magnitude of the information entropy corresponding to the audio frames to generate the audio frame sorting queue, and the audio frames with large information entropy in the audio frame sorting queue are preferentially used as the sorting targets;

[0065] The logic of the audio frame sorting operation in the sorting module is expressed as:

[0066]

[0067] In the formula: Q t is the judgment value of whether the audio frame t is a sorting target; L t is the text information converted from the audio frame based on the transliteration module; T0 is the text information converted from the audio based on the transliteration module;

[0068] Among them, based on the above formula, the judgment value of whether each audio frame in the audio frame sorting queue is a sorting target is obtained;

[0069] If the judgment value of whether the audio frame corresponds to a sorting target is 1, then the audio frame is the sorting target audio frame of the sorting module, and the sorting target audio frame is deleted by the sorting module;

[0070] Among them, when identifying the information entropy of the audio frame, with the peak frequency point in the audio frame as the boundary, the length of the shortest line segment among the line segments on both sides of the peak frequency point in the audio frame is used as the truncation reference for the other line segment, and with the peak frequency point as the starting point, a line segment equal in length to the shortest line segment is intercepted on the other line segment. Further, the segmentation accuracy is set, and based on the segmentation accuracy, the shortest line segment and the intercepted line segment among the line segments on both sides of the peak frequency point in the audio frame are segmented to obtain several groups of sub-line segments. The maximum value of the amplitude values represented at both ends of the several groups of sub-line segments is recorded as the eigenvalue in the audio frame, and the weight ε of the j-th group of eigenvalues j obeys that the closer the sub-line segment from which the eigenvalue comes is to the peak frequency point in the audio frame, the weight ε j takes a larger value, and the weight ε j > 0, and

[0071] The interaction layer includes a picking module, a recombination module, and a refreshing module. The picking module is used to pick up the audio frames remaining after sorting by the sorting module in the sorting layer, further obtain the corresponding time domain of each remaining audio frame in the audio spectrum, and intercept the audio with the same time domain in the audio based on the obtained time domain, which is denoted as sub-audio. The recombination module is used to receive all the sub-audios intercepted by the picking module, sort each group of sub-audios based on their corresponding time domains to form a new audio. The refreshing module is used to receive the new audio obtained by the operation of the recombination module, use the new audio as the audio to be transmitted by the intercom audio sender, transmit it to the intercom audio receiver, and refresh the system operation;

[0072] The transliteration module is interconnected with the segmentation module and the sorting module through local area network interaction. The transliteration module is interconnected with the marking module through local area network interaction. The marking module is interconnected with the recognition module and the receiving module through local area network interaction. The sorting module is interconnected with the picking module through local area network interaction. The picking module is interconnected with the recombination module and the refreshing module through local area network interaction.

[0073] In this embodiment, the control terminal controls the system operation. The receiving module runs to receive the audio that appears at the intercom audio sender, and synchronously converts the received audio into an audio spectrum. The recognition module synchronously obtains the audio spectrum obtained by converting the audio in the receiving module, and identifies the audio source direction based on the audio spectrum. The marking module runs later to receive the audio source direction recognition result in the recognition module, and marks the audio belonging to the audio source direction and its corresponding audio spectrum to distinguish it from other audios received by the receiving module and their corresponding audio spectra. The transliteration module further obtains the audio and its corresponding audio spectrum received in the receiving layer, further selects the audio and audio spectrum with marks, converts the selected audio into text information, and forwards the obtained audio spectrum to the segmentation module. Then the segmentation module receives the audio spectrum and performs segmentation processing on the audio spectrum. The sorting module receives in real time the audio frames obtained by segmentation in the segmentation module and the text information conversion result of the audio, sorts the audio frames in combination with the audio frames and the text information conversion result, and finally picks up the audio frames remaining after sorting by the sorting module in the sorting layer through the picking module, further obtains the corresponding time domain of each remaining audio frame in the audio spectrum, and intercepts the audio with the same time domain in the audio based on the obtained time domain, which is denoted as sub-audio. The recombination module receives in real time all the sub-audios intercepted by the picking module, sorts each group of sub-audios based on their corresponding time domains to form a new audio, and receives the new audio obtained by the operation of the recombination module through the refreshing module, uses the new audio as the audio to be transmitted by the intercom audio sender, transmits it to the intercom audio receiver, and refreshes the system operation;

[0074] Through the operation of the system in the above embodiments, when the walkie-talkie is used by the user to transmit audio, the audio is subjected to audio spectrum conversion, so as to provide audio segmentation conditions at one time. Finally, based on the segmented audio spectrum, the audio is sorted, achieving the purpose of audio streamlining and clarity, and bringing a better communication experience to walkie-talkie users.

[0075] See Figure 2 As shown, this figure further shows the poses of the three groups of microphones used in the integration of the receiving module. Based on the three groups of circular shaded areas shown in the figure, they are used to represent the microphone deployment positions;

[0076] See Figure 3 、 Figure 4 As shown, Figure 3 The example shows the audio spectrum obtained by audio conversion;

[0077] Based on Figure 4 the arrow indication in:

[0078] The vertical coordinates in the figure indicate (a); (b); (c); (d) respectively representing: Figure 3 the line spectrum pattern in the audio spectrum; several groups of audio frames obtained by segmentation; the dotted line represents the sorted target audio frames, and the solid line represents the sorted remaining audio frames; the audio spectrum reorganized after deleting the sorted target audio frames;

[0079] Based on the horizontal arrow indication and the curve arrow indication in the figure, it represents the recognition process of the information entropy of the audio frame:

[0080] Based on the curve arrow indication, taking a group in the sorted remaining audio frames as an example of the information entropy recognition target, for the convenience of display, this group of audio frames is enlarged. Further, taking the peak frequency point in the audio frame as the boundary, taking the length of the shortest line segment among the line segments on both sides of the peak frequency point in the audio frame as the truncation reference for the other line segment, and taking the peak frequency point as the starting point, a line segment with the same length as the shortest line segment is intercepted on the other line segment. The result is as shown in (1). Further, the segmentation accuracy is set. Based on the segmentation accuracy, the shortest line segment and the intercepted line segment among the line segments on both sides of the peak frequency point in the audio frame are segmented to obtain several groups of sub-line segments. The result is as shown in (2).

[0081] Embodiment 2:

[0082] At the specific implementation level, on the basis of Embodiment 1, this embodiment further specifically describes a communication information sorting system based on the analysis of large model historical parameters in Embodiment 1 with reference to Figure 1 :

[0083] The audio spectrum analyzer is deployed inside the walkie-talkie. After the microphone receives the audio sent by the user at the audio transmission end of the walkie-talkie, it transmits the audio to the audio spectrum analyzer in real time. Each group of microphones in the receiving module receives a group of audio, and each group of audio outputs its corresponding audio spectrum through the audio spectrum analyzer;

[0084] The audio source direction recognition logic in the recognition module is expressed as:

[0085]

[0086] In the formula: SIMM(a,b) is the similarity between audio spectrum a and audio spectrum b; f p1 is the value of the frequency point with the maximum energy in audio spectrum a; f p2 is the value of the frequency point with the maximum energy in audio spectrum b; f max is the value of the frequency point with the maximum energy in the two groups of audio spectra; ω1 is the weight; N is the total number of frequency bands obtained by dividing the audio spectrum; E i-a is the energy proportion of the i-th frequency band of audio spectrum a; E i-b is the energy proportion of the i-th frequency band of audio spectrum b;

[0087] Among them, based on the above formula, the similarities between three groups of audio spectra are calculated, and the f in the calculation result of the lowest similarity among the three groups of similarities is selected max as the source audio spectrum, which is used as a reference for audio source direction recognition. The microphone corresponding to the source of the audio source direction recognition reference is the audio source direction;

[0088] The weight ω1 ≤ 0.5. Before calculating the similarity between audio spectrum a and audio spectrum b, the two groups of audio spectra are divided so that the time of each sub-audio spectrum obtained by division is one second;

[0089] Before calculating the similarity between audio spectrum a and audio spectrum b, the two groups of audio spectra are normalized so that the total energy of the two groups of audio spectra is 1;

[0090] The marking content for the audio spectrum and audio corresponding to the audio source direction recognition result in the marking module is a text mark, and the text mark content is: origin, and other audio and corresponding audio spectra, that is, the remaining two groups of audio and audio spectra.

[0091] In this embodiment, through the above settings, the audio source direction recognition logic in the recognition module is further defined, and based on this as historical parameters, the usage habits of walkie-talkie users are analyzed. When the user operates the walkie-talkie, the audio source direction is analyzed, so that the audio captured by the three groups of microphones integrated in the receiving module can be more appropriately selected by the system, thereby improving the operation accuracy of the system.

[0092] Embodiment 3:

[0093] At the specific implementation level, based on Embodiment 1, this embodiment refers to Figure 1 to further specifically describe a communication information sorting system based on the analysis of large model historical parameters in Embodiment 1:

[0094] After receiving the audio spectrum, the segmentation module picks up all the valley and peak frequency points in the audio spectrum, and takes the spectral segments where every three adjacent groups of frequency points are located as a group of audio frames. Among the three groups of adjacent frequency points, there are two groups of valley frequency points and one group of peak frequency points, and the peak frequency point is between the two groups of valley frequency points. The above is denoted as the audio frame determination logic. The segmentation module segments the audio spectrum based on the audio frame determination logic to obtain several groups of audio frames, and the unit time of several groups of audio frames is not equal;

[0095] After the segmentation module segments the audio spectrum to obtain several groups of audio frames, the segmentation module synchronously obtains the audio obtained in the transliteration module, and based on the ratio of the respective unit times of several groups of audio frames, performs segmentation processing on the audio to obtain several groups of sub-audios. After several groups of sub-audios are configured behind the corresponding audio frames, they are sent to the transliteration module to perform text information conversion operations on each group of sub-audios, and further bind the converted text information to the audio frames corresponding to the sub-audios from which the text information is sourced, and then return to the segmentation module. Based on the segmentation module, the mutually bound text information and audio frames are sent to the sorting module.

[0096] Through the above settings, the operation logic of the system in Embodiment 1 for audio frame segmentation of the audio spectrum is defined, ensuring the stable operation of the segmentation module in the system to output audio frames, and providing further operation data support for the operation of the analysis module and the system interaction layer.

[0097] As Figure 1 shown, when the recombination module performs sorting and recombination operations on each group of sub-audios, fade-in and fade-out processing is performed on two adjacent groups of sub-audios;

[0098] Among them, when fade-in and fade-out processing is performed on two adjacent groups of sub-audios, the fade-in and fade-out duration follows: the longer the duration of the sub-audio with the shortest time domain in two adjacent groups of sub-audios, the longer the fade-in and fade-out duration, and vice versa, the shorter the fade-in and fade-out duration; when fade-in and fade-out processing is performed on two adjacent groups of sub-audios, the fade-in and fade-out audio intensity is within the audio intensity threshold composed of the maximum audio intensity and the minimum audio intensity in two adjacent groups of sub-audios;

[0099] During the operation stage of the refresh module, the recognition result of the audio source direction recognized by the recognition module in the receiving layer during the current operation of the system is synchronously recorded. When the audio source directions recorded twice in a row are the same, the system controls the recognition module not to run during the next operation based on the refresh module, and applies the previous audio source direction recognition result to the marking module;

[0100] Among them, during each system operation process, the system operation process in which the recognition module does not operate is not used as the determination target for the audio source directions of two consecutive recordings to be the same.

[0101] Through the above settings, a further continuous operation logic is provided for the system in Embodiment 1, ensuring that the system can stably and continuously serve the walkie-talkie communication.

[0102] In summary, during the operation of the system in the above embodiments, by setting multiple groups of microphones, the historical usage habits of walkie-talkie users are collected. Then, based on the historical usage habits of walkie-talkie users, the user audio received by multiple groups of microphones is selected to ensure that the system can use the group of audio that is the clearest and closest to the original meaning of the walkie-talkie user audio as the further operation target of the system. Thereby, operations such as splitting the corresponding audio spectrum of the audio and converting text information are performed to achieve the ultimate goal of eliminating the audio frames that do not contain information in the audio. Finally, the audio transmitted by the walkie-talkie is made clearer, more concise, and contains less noise. Furthermore, when the user uses the walkie-talkie, a better communication experience is brought to the user based on this system.

[0103] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A communication information sorting system based on the analysis of historical parameters of a large model, characterized in that It includes a control terminal, a receiving layer, a sorting layer, and an interaction layer; The control terminal is the main control end of the system and is used to decide the start and stop of the system and the start and stop of each module in the system; The audio that appears at the walkie-talkie audio sending end is received by the receiving layer. At the same time, the corresponding audio spectrum is obtained based on the received audio. The direction of the audio source is identified through the audio spectrum. The audio and the audio spectrum are marked based on the identified audio source direction. The sorting layer receives the audio and the audio spectrum, and uses the audio spectrum to perform segmentation processing on the audio to obtain several groups of audio frames corresponding to different unit times. Further, the sorted remaining audio frames are fed back to the interaction layer. The interaction layer recombines the received audio frames and sends the combined audio to the walkie-talkie audio receiving end; The sorting layer includes a transliteration module, a segmentation module, and a sorting module. The transliteration module is used to obtain the audio and the corresponding audio spectrum received in the receiving layer, further select the audio and the audio spectrum with marks, convert the selected audio into text information, and forward the obtained audio spectrum to the segmentation module. The segmentation module is used to receive the audio spectrum and perform segmentation processing on the audio spectrum. The sorting module is used to receive the audio frames obtained by segmentation in the segmentation module and the text information conversion result of the audio, and sort the audio frames in combination with the audio frames and the text information conversion result; When the sorting module receives the audio frames and the text information conversion result of the audio, it synchronously performs information entropy recognition on each audio frame. Based on the information entropy recognition result of the audio frame, an audio frame sorting queue is generated, and then the audio frame sorting operation is executed; The recognition logic of the information entropy of the audio frame is expressed as: ; Wherein: is the information entropy of the audio frame; is the total amount of eigenvalue in the audio frame; is the probability that the j-th group of eigenvalues appears; is the weight of the j-th group of eigenvalues; is the correction; Among them, when identifying the information entropy of an audio frame, taking the peak frequency point in the audio frame as the boundary, using the length of the shortest line segment among the line segments on both sides of the peak frequency point in the audio frame as the truncation reference for the other line segment, taking the peak frequency point as the starting point, intercepting a line segment with the same length as the shortest line segment on the other line segment, further setting the segmentation accuracy, and based on the segmentation accuracy, segmenting the shortest line segment and the intercepted line segment among the line segments on both sides of the peak frequency point in the audio frame to obtain several groups of sub-line segments. The maximum value of the amplitude values represented by both ends of the several groups of sub-line segments is denoted as the eigenvalue in the audio frame, and the weight of the jth group of eigenvalues obeys that the closer the sub-line segment from which the eigenvalue is derived is to the peak frequency point in the audio frame, the larger the weight value is, and the weight > 0, and .

2. The communication information sorting system based on the analysis of the historical parameters of the large model according to claim 1, wherein The receiving layer includes a receiving module, an identification module, and a marking module. The receiving module is used to receive the audio that appears at the walkie-talkie audio sending end and synchronously convert the received audio into an audio spectrum. The identification module is used to obtain the audio spectrum obtained by converting the audio in the receiving module and identify the direction of the audio source based on the audio spectrum. The marking module is used to receive the audio source direction identification result in the identification module and mark the audio belonging to the audio source direction and its corresponding audio spectrum to distinguish it from other audio and the corresponding audio spectrum received by the receiving module; Among them, the receiving module is integrated by three groups of microphones. The three groups of microphones are distributed in a triangular shape, and the distance between each adjacent two groups of microphones is equal. The receiving module is installed on the walkie-talkie, receives the audio sent by the user at the walkie-talkie audio sending end in real time, and imports the audio spectrum into the audio spectrum analyzer, and outputs the audio spectrum corresponding to the audio through the audio spectrum analyzer.

3. A communication information sorting system based on the analysis of large model historical parameters according to claim 2, characterized in that, The audio spectrum analyzer is deployed inside the walkie-talkie. After the microphone receives the audio sent by the user at the walkie-talkie audio sending end, it is transmitted to the audio spectrum analyzer in real time. Each group of microphones in the receiving module receives a group of audio, and each group of audio outputs its corresponding audio spectrum through the audio spectrum analyzer; The audio source direction identification logic in the identification module is expressed as: ; Wherein: is the similarity between audio frequency spectrum a and audio frequency spectrum b; is the value of the frequency point with the maximum energy in audio frequency spectrum a; is the value of the frequency point with the maximum energy in audio frequency spectrum b; is the value of the frequency point with the maximum energy in the two groups of audio frequency spectra; is the weight; is the total number of frequency bands obtained by dividing the audio frequency spectrum; is the energy proportion of the i-th frequency band of audio frequency spectrum a; is the energy proportion of the i-th frequency band of audio frequency spectrum b; Among them, the similarity between the three groups of audio spectra is calculated based on the above formula, and the audio spectrum with the lowest similarity in the calculation results of a group of similarities is selected as the reference for identifying the audio source direction. The microphone from which the reference for identifying the audio source direction is derived is the audio source direction.

4. A communication information sorting system based on the analysis of historical parameters of a large model according to claim 2 or 3, characterized in that Weight ≤0.

5. Before performing the similarity calculation on the audio spectrum a and the audio spectrum b, the two sets of audio spectra are divided so that the time of each sub-audio spectrum obtained by the division is one second; Before calculating the similarity between the audio spectrum a and the audio spectrum b, the two groups of audio spectra are normalized so that the total energy of the two groups of audio spectra is 1; The marked content for the audio spectrum and audio corresponding to the recognition result of the audio source direction in the marking module is a text mark, and the text mark content is: origin, and the other audio and corresponding audio spectrum, that is, the remaining two groups of audio and audio spectrum.

5. A communication information sorting system based on the analysis of historical parameters of a large model according to claim 1, characterized in that, After receiving the audio spectrum, the segmentation module picks up all the valley-peak frequency points in the audio spectrum, and takes the spectrum segment where every three adjacent frequency points are located as a group of audio frames. Among the three adjacent frequency points, there are two groups of valley frequency points and one group of peak frequency points, and the peak frequency point is between the two groups of valley frequency points. The above is recorded as the audio frame determination logic. The segmentation module segments the audio spectrum based on the audio frame determination logic to obtain several groups of audio frames, and the unit time of several groups of the audio frames is not equal; After the segmentation module segments the audio spectrum to obtain several groups of audio frames, the segmentation module synchronously obtains the audio obtained in the transliteration module, and based on the ratio of the respective unit times of several groups of audio frames, performs segmentation processing on the audio to obtain several groups of sub-audios. After several groups of sub-audios are configured behind the corresponding audio frames, they are sent to the transliteration module to perform text information conversion operations on each group of sub-audios, and further bind the converted text information to the audio frames corresponding to the sub-audios where the text information source is located, and then return to the segmentation module. Based on the segmentation module, the mutually bound text information and audio frames are sent to the sorting module.

6. The communication information sorting system based on the analysis of the historical parameters of the large model according to claim 1, characterized in that, 2 > ≥ 1, and follows: the greater the length difference between the line segments on both sides of the peak frequency point in the audio frame, the greater the correction value, and vice versa, the smaller the correction value; When the sorting module receives the audio frames and the text information conversion results of the audio, that is, the mutually bound text information and audio frames, when the sorting module runs to generate an audio frame sorting queue, it sorts each audio frame in descending order according to the magnitude of the information entropy corresponding to the audio frame to generate an audio frame sorting queue, and gives priority to the audio frames with large information entropy in the audio frame sorting queue as the sorting targets; The logic representation of the audio frame sorting operation in the sorting module is: ; Wherein: is the determination value of whether the audio frame t is a sorting target; is the text information converted from the audio frame based on the transliteration module; is the text information converted from the audio based on the transliteration module; Among them, based on the above formula, the determination value of whether each audio frame in the audio frame sorting queue is a sorting target is obtained; If the determination value of whether the audio frame corresponds to a sorting target is 1, then the audio frame is a sorting target audio frame of the sorting module, and the sorting target audio frame is deleted by the sorting module.

7. A communication information sorting system based on the analysis of historical parameters of a large model according to claim 2, characterized in that, The interaction layer includes a picking module, a recombination module, and a refreshing module. The picking module is used to pick up the audio frames remaining after sorting by the sorting module in the sorting layer, and further obtain the time domain corresponding to each remaining sorted audio frame in the audio spectrum, and intercept the audio with the same time domain in the audio according to the obtained time domain, which is recorded as a sub-audio. The recombination module is used to receive all the sub-audios intercepted by the picking module, sort each group of sub-audios based on their corresponding time domains to form a new audio. The refreshing module is used to receive the new audio obtained by the operation of the recombination module, use the new audio as the audio to be transmitted by the audio sending end of the walkie-talkie, transmit it to the audio receiving end of the walkie-talkie, and refresh the system operation.

8. A communication information sorting system based on the analysis of the historical parameters of a large model according to claim 7, characterized in that, When the recombination module performs sorting and recombination operations on each group of sub-audios, fade-in and fade-out processing is performed on each adjacent two groups of sub-audios; Among them, when fade-in and fade-out processing is performed on two adjacent groups of sub-audios, the fade-in and fade-out duration follows: the longer the duration of the sub-audio with the shortest time domain in two adjacent groups of sub-audios, the longer the fade-in and fade-out duration; conversely, the shorter the fade-in and fade-out duration. When fade-in and fade-out processing is performed on two adjacent groups of sub-audios, the fade-in and fade-out audio intensity is within the audio intensity threshold composed of the maximum audio intensity and the minimum audio intensity in two adjacent groups of sub-audios.

9. A communication information sorting system based on the analysis of historical parameters of a large model according to claim 7, characterized in that, During the operation stage of the refresh module, the recognition result of the audio source direction recognized by the recognition module in the receiving layer during the current operation of the system is synchronously recorded. When the audio source directions recorded twice in a row are the same, the system controls the recognition module not to operate during the next operation based on the refresh module, and applies the previous audio source direction recognition result to the marking module. Among them, during each operation of the system, the operation process of the system in which the recognition module does not operate is not used as the determination target for the audio source directions recorded twice in a row to be the same.

10. A communication information sorting system based on the analysis of large model historical parameters according to claim 7, characterized in that, The transliteration module is interconnected with a segmentation module and a sorting module through local area network interaction. The transliteration module is interconnected with a marking module through local area network interaction. The marking module is interconnected with a recognition module and a receiving module through local area network interaction. The sorting module is interconnected with a picking module through local area network interaction. The picking module is interconnected with a recombination module and a refresh module through local area network interaction.

Citation Information

Patent Citations

  • Communication method and system of interphone

    CN117119553A

  • Voice endpoint detection method and device, equipment and storage medium

    CN113314153A

  • Audio data transmission system of communication equipment

    CN117255383A