A voice interaction optimization method and device based on cross-positioning
Through the combination of cross-position technology of multiple sets of sound signal acquisition nodes, combined with audio signal processing and cloud database, the problem of inaccurate voice signal positioning in scenes of multiple people simultaneously is solved, and efficient voice interaction in complex environments is achieved.
Patent Information
- Application Number
- CN202510764739.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing voice interaction technology is difficult to accurately distinguish the spatial characteristics of each voice signal in complex scenarios where multiple people speak at the same time, resulting in low interaction accuracy and efficiency.
Multiple groups of different sound signal acquisition nodes are combined, and by determining the selected surface and finding the intersection lines formed by any two selected surfaces, the spatial characteristics of the relevant voice signals are determined, combined with the audio signal waveform and frequency characteristics for audio processing, abnormal signals are eliminated, and text output is output using a self-built cloud database.
It improves the reliability and accuracy of voice signal positioning, removes noise interference, ensures accurate identification of voice signal sources in complex environments, and improves the accuracy and interaction efficiency of text conversion.
Smart Images

Figure CN120279900B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voice interaction and provides a voice interaction optimization method and device based on cross-positioning. Background Art
[0002] With the rapid development of artificial intelligence (AI), voice interaction, as a natural and convenient method of human-computer interaction, is being widely used in a wide range of fields, including smart homes, intelligent customer service, in-vehicle systems, and smart wearable devices. However, in practical application scenarios, voice interaction faces many challenges, especially in complex cross-scenario environments, where existing voice interaction technologies have certain limitations.
[0003] Traditional voice interaction systems often lack the ability to accurately identify the spatial location of voice signals. Most systems simply receive voice input without being able to clearly identify the direction and specific location of the voice signal. In complex scenarios where multiple people speak simultaneously, such as conference rooms and restaurants, voice signals from different locations intersect. Existing technologies struggle to accurately distinguish the spatial characteristics of each voice signal, making it difficult to effectively process the voices of different speakers, impacting the accuracy and efficiency of interactions.
[0004] Therefore, it is necessary to provide a voice interaction optimization method based on cross-localization to solve the above problems. Summary of the Invention
[0005] The present invention provides a voice interaction optimization method and device based on cross-localization to solve the technical problems in the prior art that the source direction and specific location of the voice signal cannot be clearly determined, especially in complex scenarios where multiple people speak at the same time (such as conference rooms, restaurants, etc. where voice signals from different locations are intertwined), it is difficult to accurately distinguish the spatial characteristics of each voice signal, resulting in the inability to effectively perform targeted processing on the voices of different people, affecting the accuracy and efficiency of the interaction. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0006] The first aspect of the present invention proposes a voice interaction optimization method based on cross-positioning, which includes: collecting voice signals from at least three collection nodes, confirming the spatial position information associated with different voice signals, so as to obtain the spatial features associated with each voice signal, specifically including: confirming the time difference corresponding to two groups of associated moments in a pairwise manner, determining the feature surface, and constructing multiple vertical surfaces perpendicular to the feature surface; based on the feature surface and multiple vertical surfaces, confirming the feature difference to determine the selected surface; using the intersection line formed by any two selected surfaces as the spatial feature of the relevant voice signal; based on the obtained spatial features associated with each voice signal, performing audio processing on multiple groups of voice signals belonging to the same spatial feature, and then performing text output on the voice signals belonging to the same spatial feature to confirm the feature output text; based on a self-built cloud database, extracting the storage content associated with each feature output text according to the feature output text generated by each spatial feature, and compiling each feature output text into an index text; when receiving user input, returning the associated request data content based on the index text recognized from the user input.
[0007] The second aspect of the present invention proposes a voice interaction optimization device based on cross-positioning, which executes the voice interaction optimization method based on cross-positioning described in the first aspect of the present invention. The voice interaction optimization device includes: an acquisition and processing module for acquiring voice signals from at least three acquisition nodes, confirming the spatial position information associated with different voice signals, so as to obtain the spatial features associated with each voice signal, specifically including: confirming the time difference corresponding to the two groups of associated moments in a two-by-two manner, determining the feature surface, and constructing multiple vertical surfaces perpendicular to the feature surface; based on the feature surface and multiple vertical surfaces, confirming the feature difference to determine the selected surface; combining any two selected surfaces The formed cross lines are used as spatial features of related voice signals; the confirmation module, based on the spatial features associated with each voice signal obtained, performs audio processing on multiple groups of voice signals belonging to the same spatial feature, and then outputs text for the voice signals belonging to the same spatial feature to confirm the feature output text; the extraction and processing module, based on a self-built cloud database, extracts the storage content associated with each feature output text according to the feature output text generated by each spatial feature, and compiles each feature output text into an index text; the query determination module, when receiving user input, returns the associated request data content based on the index text identified from the user input.
[0008] The third aspect of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the cross-localization-based voice interaction optimization method described in the first aspect of the present invention.
[0009] A fourth aspect of the present invention provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the cross-localization-based voice interaction optimization method described in the first aspect of the present invention.
[0010] The embodiments of the present invention include the following advantages:
[0011] Compared with the prior art, the present invention uses a combination of multiple groups of different sound signal acquisition nodes for processing. The spatial characteristics of the relevant voice signals are determined by determining different selected surfaces and finding the intersection lines formed by any two selected surfaces. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces positioning errors, and ensures that the source position of each voice signal can be accurately identified even in a complex multi-voice environment. According to the spatial characteristics of the voice signal, audio processing is performed on the voice signal in the same space, and the audio characteristics of the voice signal are determined by calculating the peak point, horizontal and vertical distance, and audio frequency of the audio signal waveform. Then, the continuously generated voice signal is averaged and analyzed for fluctuation intervals to eliminate voice signals with abnormal characteristics, effectively remove interference signals such as noise, and improve the quality and purity of the voice signal, making the subsequent text conversion more accurate and reliable. It can not only fully locate, but also effectively interact and remove noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of an example of a cross-localization-based voice interaction optimization method of the present invention;
[0013] Figure 2 This is a schematic diagram showing the principle of the cross-localization-based voice interaction optimization method of the present invention from another perspective;
[0014] Figure 3 It is a structural block diagram of the cross-positioning-based voice interaction optimization device of the present invention;
[0015] Figure 4 is a schematic structural diagram of an electronic device according to an embodiment of the present invention;
[0016] Figure 5 is a schematic structural diagram of an embodiment of a computer-readable medium according to the present invention. DETAILED DESCRIPTION
[0017] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0018] Traditional voice interaction systems often lack the ability to accurately identify the spatial location of voice signals. They can simply receive voice input without being able to pinpoint the direction and specific location of the voice signal. In complex scenarios where multiple people speak simultaneously, such as conference rooms and restaurants, voice signals from different locations intersect. Existing technologies struggle to accurately distinguish the spatial characteristics of each voice signal, making it difficult to effectively process the voices of different speakers, impacting the accuracy and efficiency of interaction.
[0019] In view of the above problems, the present invention proposes a voice interaction optimization method based on cross-positioning, which uses a combination of multiple groups of different sound signal acquisition nodes for processing. The spatial characteristics of the relevant voice signals are determined by determining different selected surfaces and finding the intersection line formed by any two selected surfaces. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces positioning errors, and ensures that the source position of each voice signal can be accurately identified even in a complex multi-voice environment. According to the spatial characteristics of the voice signal, audio processing is performed on the voice signal in the same space, and the audio characteristics of the voice signal are determined by calculating the peak point, horizontal and vertical distance, and audio frequency of the audio signal waveform. Then, the continuously generated voice signal is averaged and analyzed for fluctuation ranges to eliminate voice signals with abnormal characteristics, effectively remove interference signals such as noise, improve the quality and purity of the voice signal, and make the subsequent text conversion more accurate and reliable. It can not only fully locate, but also effectively interact and remove noise.
[0020] Example 1
[0021] Refer to the following Figure 1 、 Figure 2 , the contents of the present invention will be described in detail.
[0022] Figure 1 This is a flowchart of an example of the cross-localization-based voice interaction optimization method of the present invention.
[0023] like Figure 1 As shown, in step S101, speech signals from at least three acquisition nodes are collected, and spatial position information associated with different speech signals is confirmed to obtain spatial features associated with each speech signal.
[0024] Specifically, voice signals are collected from at least three collection nodes. For example, there may be three collection nodes, specifically including a first collection node (e.g., a voice interaction device, specifically including a collector, a smart speaker, a car speaker, or a wearable device), and second and third collection nodes disposed on the top and / or side of the voice interaction device. As long as the second and third collection nodes are asymmetrically disposed, and the three groups of sound signal collection points of the first, second, and third collection nodes are non-symmetrically disposed, thereby forming a staggered or crossed pattern, i.e., a staggered or staggered arrangement, this allows for more effective feature locking and facilitates subsequent specific confirmation of spatial location.
[0025] According to the collected voice signals, the spatial position information associated with different voice signals is determined, specifically comprising the following steps:
[0026] Step S201: confirm the associated moments of the voice signals received from different acquisition nodes, confirm the time difference corresponding to the two groups of associated moments in a pairwise manner, use any one group of associated moments as the pre-feature, and use the other group of associated moments as the post-feature, use the two acquisition nodes associated with the time difference between the pre-feature and the post-feature as the first feature point and the second feature point, determine the feature surface based on the first feature point and the second feature point, and construct multiple vertical surfaces perpendicular to the feature surface.
[0027] Specifically, for a speech signal, there are three collected associated time instants t1, t2, and t3, and there are time differences |t1-t2|, |t2-t3|, and |t1-t3| between each time instant.
[0028] The associated moments at which different sound signal collection nodes receive corresponding voice signals are identified and labeled as Si, where i represents a different sound signal collection node, i.e., the i-th collection node, and i is a positive integer, specifically 1, 2, or 3 in the present invention. The time difference between the corresponding associated moments of two groups of different sound signal collection nodes is identified in pairs. The associated moments of one group are used as the pre-feature T1, and the associated moments of the other group are used as the post-feature T2. The time difference value is calculated as T1-T2 to identify the associated time difference between the corresponding sound signal collection nodes.
[0029] Next, two collection nodes of two different sound signals associated with the time difference, specifically the first collection node and the second collection node, are used as the first feature point and the second feature point.
[0030] For example, Figure 2Midpoints A, B, and C are used as the first, second, and third feature points, respectively. A feature surface is determined based on the first and second feature points. Specifically, point A and point B are connected to form a line segment AB passing through points A and B. The plane corresponding to line segment AB is used as the feature surface, and multiple perpendicular planes perpendicular to the feature surface are constructed. Point A, point B, and point C correspond to the first, second, and third acquisition points, respectively.
[0031] Step S202: Based on the feature surface and multiple vertical surfaces, select a point to be used as the sound point to be confirmed, identify the distance between the point to be sounded and the first feature point, and the distance between the point to be sounded and the second feature point, so as to calculate the first time feature value of the acquisition node corresponding to the first feature point and the second time feature value of the acquisition node corresponding to the second feature point, and further confirm the feature difference to determine the selected surface.
[0032] A group of points are randomly selected from different vertical planes and recorded as the sound points to be confirmed. The straight-line distances L1 and L2 between the sound points to be confirmed and the first feature point and the second feature point are identified respectively. Then, L1÷c=t1 and L2÷c=t2 are used to confirm the first time characteristic value t1 and the second time characteristic value t2 associated with the first acquisition node and the second acquisition node of the two groups of different sound signals, where c is the preset sound propagation speed in the air, for example, 343m / s.
[0033] Based on the confirmed time feature values t1 and t2, the time feature value associated with the front feature T1 is placed in front, and the time feature value associated with the back feature T2 is placed in the back, and difference processing is performed to confirm the feature difference. From multiple vertical planes, select or lock the plane with the same feature difference and time difference, that is, the first selected plane (for example Figure 2 ).
[0034] Step S203: Repeat the relevant steps S201 and S202 (ie, the calculation and determination steps) to obtain a second selected surface determined based on the first acquisition node and the third acquisition node, and a third selected surface determined based on the second acquisition node and the third acquisition node.
[0035] It should be noted that the corresponding acquisition device has three acquisition nodes, and between each of the two acquisition nodes lies a plane, i.e., a corresponding feature plane. Several vertical planes can be identified on the surface of the feature plane. When the time signature associated with a corresponding point within the vertical plane matches the time signature associated with the third acquisition node, that is, the location of the selected plane, the sound originating point is located within the corresponding selected plane. The above is provided as an optional example only and should not be construed as limiting the present invention.
[0036] Step S204: using an intersection line formed by any two selected surfaces among the first selected surface, the second selected surface, and the third selected surface as a spatial feature of the relevant speech signal.
[0037] Specifically, the intersection line generated by any two selected planes in space is identified (since the selected planes cannot be parallel, an intersection line must exist), and the spatial location information of the resulting intersection line is used as the spatial feature of the relevant speech signal. Specifically, the spatial location information includes the length of the intersection line, the angle formed by the intersection line and the vertical direction (or the horizontal direction perpendicular to the vertical direction), and the distance between the intersection line and a specified point (such as the first acquisition node, the second acquisition node, the third acquisition node, or other points).
[0038] By confirming the spatial position information associated with different voice signals, the spatial features associated with each voice signal can be obtained. Specifically, the intersection lines generated by the selected surface are determined as the spatial features associated with the relevant voice signals. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces positioning errors, and ensures that the source position of each voice signal can be accurately identified even in complex multi-voice environments.
[0039] It should be noted that the above is merely provided as an optional example and should not be construed as a limitation to the present invention.
[0040] Next, in step S102, based on the spatial features associated with each speech signal obtained, multiple groups of speech signals belonging to the same spatial feature are audio processed, and then text output is performed on the speech signals belonging to the same spatial feature to confirm the feature output text.
[0041] In a specific embodiment, the voice signals a and the voice signals b belonging to the same spatial feature (e.g., the cross line x) are audio-processed, and related voice signals with inconsistent audio features are preliminarily eliminated. The voice signals a and the voice signals b are then audio-processed to output feature text, and the feature output text is confirmed.
[0042] When multiple speech signals correspond to the same spatial feature, the audio features associated with each different speech signal are different. Different speech signals can be distinguished by confirming the audio features.
[0043] To confirm the feature output text, the following steps are specifically included:
[0044] Step S301: receiving multiple groups of speech signals belonging to the same spatial feature, and confirming the audio features associated with the corresponding speech signals based on the audio signal waveforms associated with the different speech signals.
[0045] For example, a speech signal a and a speech signal b belonging to the same spatial feature (eg, a cross line x) are received, and based on the audio signal waveforms associated with the speech signal a and the speech signal b.
[0046] Based on the audio signal waveforms associated with the voice signal a and the voice signal b, for a single group of voice signals (such as voice signal a or voice signal b), the peak points in the current audio signal waveform are confirmed from the current audio signal waveform, and the horizontal and vertical distances between adjacent peak points are recorded and represented by Fk, where k represents different adjacent peak points. The audio frequencies associated with each group of peak points are then confirmed and represented by Pk. Several groups of horizontal and vertical distances Fk confirmed in the current audio signal waveform are averaged to confirm the distance mean JJ. Then, several groups of audio frequencies Pk confirmed are averaged to confirm the audio mean YY of the current audio signal.
[0047] The audio features of the speech signal are calculated using the following expressions:
[0048] TT=JJ×C1+YY×C2
[0049] Among them, TT represents the audio characteristics of the speech signal; JJ represents the distance mean confirmed by averaging several groups of horizontal and vertical distances confirmed in the current audio signal waveform; YY represents the audio mean of the current audio signal confirmed by averaging several groups of audio frequencies; C1 and C2 are the first fixed coefficient factor and the second fixed coefficient factor respectively, the value of C1 is 0.5~0.7, which is 0.574 in this example, and the value of C2 is 0.4~0.5, which is 0.426 in this example.
[0050] Step S302: performing audio processing on multiple groups of speech signals continuously generated by the same spatial feature, confirming the audio features associated with different speech signals, and reprocessing the multiple groups of audio features TT.
[0051] The so-called continuity means that the time difference between different voice signals is less than 1 second, and this time difference represents the continuity of the corresponding voice signals.
[0052] Several groups of audio features TT are reprocessed, specifically including averaging multiple groups of audio features, locking the audio feature mean, and confirming the fluctuation range [YJ-Y2, YJ+Y2] based on the preset value Y2, where YJ is the locked audio feature mean, Y2 is the preset value, and the value of Y2 is 2 to 4, for example, determined in advance by relevant operators based on experience, keeping the range of the above-mentioned fluctuation range unchanged, making the endpoint values of the fluctuation range change synchronously, and recording the audio features TT included in the fluctuation range corresponding to each different numerical range, recording the fluctuation range with the largest total number of audio features as the determined interval, and recording the audio features TT that do not belong to the determined interval as abnormal features, eliminating the voice signals associated with the abnormal features, and re-sorting the multiple groups of continuously generated voice signals currently detected to confirm the voice signal sorting sequence.
[0053] Next, the voice signal sorting sequence is subjected to analog-to-digital processing to lock the digital signal associated with the corresponding voice signal, and then the associated characters are locked based on the digital signal. The locked characters are kept in the original sorting method unchanged to generate the feature output text associated with the corresponding spatial feature. For example: "How is the temperature today?", "Today" corresponds to a digital signal, "Day" corresponds to a digital signal, and so on. The corresponding characters are gradually confirmed to determine the corresponding output text, thereby locking the feature output text: "How is the temperature today?"
[0054] Specifically, the corresponding voice content is generated into text content, and the relevant conversion can be performed automatically. The abnormal voice signal associated with the corresponding spatial feature is generally noise, and its signal characteristics are significantly different from other signal characteristics. Specifically, these abnormal audio such as noise are removed.
[0055] It should be noted that the above is merely provided as an optional example and should not be construed as a limitation to the present invention.
[0056] Next, in step S103, based on the self-built cloud database, according to the feature output texts generated by each spatial feature, the storage content associated with each feature output text is extracted, and each feature output text is compiled into an index text.
[0057] Specifically, the self-built cloud database includes various titles (ie, headers), and each title corresponds to a storage area for storing feature output text associated with the spatial feature.
[0058] It should be noted that, in this example, the header can be understood as the title of a storage area, and for each title, there is corresponding stored text output content.
[0059] S41. Mark the different feature output texts associated with different spatial features (i.e., the spatial features associated with each voice signal obtained in step S102) as index texts, and extract the headers of other output data (i.e., the corresponding index texts) from the cloud database. The headers are all preset in advance by relevant personnel. The index texts are compared and verified with different headers to lock the feature headers.
[0060] Specifically include the following steps.
[0061] S411: Identify the text characters in the feature output text that are identical in the corresponding header and index text, and record the number G1 of identical text characters, then record the total number G2 of text characters in the corresponding header and the total number G3 of index text, and confirm the comprehensive proportion of identical text characters.
[0062] Preferably, the following expression is used to calculate the comprehensive proportion of each text character to determine the proportion of the first group:
[0063] ZB1=G1÷G2
[0064] Among them, ZB1 represents the first group of proportion values calculated and confirmed, G1 represents the number of identical text characters recorded in the corresponding header of the recognition feature output text and the index text; G2 represents the total number of text characters recorded in the corresponding header of the recognition feature output text and the index text.
[0065] Use the following expression to calculate the comprehensive proportion of each text character and confirm the proportion of the second group:
[0066] ZB2=G1÷G3
[0067] Among them, ZB2 represents the second group of proportion values calculated and confirmed; G1 represents the number of identical text characters recorded in the recognition feature output text whose corresponding headers are the same as the index text; G3 represents the total number of index texts recorded in the recognition feature output text whose corresponding headers are the same as the index text.
[0068] Specifically, the first group of proportion values ZB1 and the second group of proportion values ZB2 are averaged to lock the comprehensive proportion associated with the same text characters.
[0069] S412: If there is no header with the same text characters, an error signal display is directly generated.
[0070] If there is a header with a group of identical text characters, the header is marked as a characteristic header.
[0071] If there are multiple headers with the same text characters, the header associated with the largest comprehensive proportion is selected from the comprehensive proportions associated with each different header and recorded as the selected header, and the selected header is used as the feature header.
[0072] For example, the proposed feature output text is "How is the temperature today?" In its cloud database, there are headers for the corresponding output data: the headers can be "Today's temperature", "Yesterday's temperature", "The day before yesterday's temperature", "City temperature", etc., or other headers.
[0073] Step S42: confirming the index path of the feature header in the feature output text, and outputting the storage content associated with the determined index path in voice to complete the real-time interaction.
[0074] For example, the proposed feature output text is: "What is the temperature today?"
[0075] The cloud database stores headers corresponding to the output data, including the following headers: "Today's Temperature", "Yesterday's Temperature", "The Day Before Yesterday's Temperature", "City Temperature", etc., and may also store other headers.
[0076] It's important to note that in this example, a header is understood as the title of a storage area. For this title, there's corresponding stored output content. If there's only a single header (that is, identical text characters), the content is displayed directly. If there are multiple headers, "Today's Temperature" is selected as the feature header and the corresponding content is displayed and output.
[0077] Based on the header and feature output text, it can be determined that "What's the temperature today" has the highest overlap with the header "Today's temperature", that is, the proportion of the first group is 4 / 7, the proportion of the second group is 1, and the confirmed comprehensive proportion is 5.5 / 7.
[0078] There is only a single header (that is, the same text characters exist), and the content can be displayed directly.
[0079] If there are multiple headers, one of them (such as "today's temperature") is selected as the feature header, and the relevant content is displayed and output.
[0080] It should be noted that the above is merely provided as an optional example and should not be construed as a limitation to the present invention.
[0081] Next, in step S104 , when user input is received, the associated requested data content is returned based on the index text recognized from the user input.
[0082] For example, when a user inputs "What's the temperature today?", the system recognizes index text (such as "today's temperature") from the user input and returns the associated requested data content: such as 25 degrees.
[0083] It should be noted that the above is merely provided as an optional example and should not be construed as a limitation to the present invention.
[0084] Compared with the prior art, the present invention uses a combination of multiple groups of different sound signal acquisition nodes for processing. The spatial characteristics of the relevant voice signals are determined by determining different selected surfaces and finding the intersection lines formed by any two selected surfaces. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces positioning errors, and ensures that the source position of each voice signal can be accurately identified even in a complex multi-voice environment. According to the spatial characteristics of the voice signal, audio processing is performed on the voice signal in the same space, and the audio characteristics of the voice signal are determined by calculating the peak point, horizontal and vertical distance, and audio frequency of the audio signal waveform. Then, the continuously generated voice signal is averaged and analyzed for fluctuation intervals to eliminate voice signals with abnormal characteristics, effectively remove interference signals such as noise, and improve the quality and purity of the voice signal, making the subsequent text conversion more accurate and reliable. It can not only fully locate, but also effectively interact and remove noise.
[0085] Example 2
[0086] The following are embodiments of the apparatus of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the apparatus embodiments of the present invention, please refer to the method embodiments of the present invention.
[0087] Figure 3 This is a schematic diagram of an example of a cross-positioning-based voice interaction optimization device according to the present invention. Figure 3 , a voice interaction optimization device is described. The voice interaction optimization device is used to execute the voice interaction optimization method described in the first aspect of the present invention.
[0088] like Figure 3 As shown, the voice interaction optimization device 300 includes a collection and processing module 310, a confirmation module 320, an extraction and processing module 330, and a query determination module 340.
[0089] In one specific embodiment, the acquisition and processing module 310 is configured to collect speech signals from at least three acquisition nodes and identify the spatial location information associated with different speech signals to obtain spatial features associated with each speech signal. Specifically, this includes: determining the time difference between two sets of associated moments in a pairwise manner, determining a characteristic surface, and constructing multiple vertical surfaces perpendicular to the characteristic surface; determining the characteristic difference based on the characteristic surface and multiple vertical surfaces to determine a selected surface; and using the intersection formed by any two selected surfaces as the spatial feature of the relevant speech signal. The confirmation module 320 performs audio processing on multiple sets of speech signals belonging to the same spatial feature based on the obtained spatial features associated with each speech signal, then outputs the speech signals belonging to the same spatial feature as text and identifies the characteristic output text. The extraction and processing module 330, based on a self-built cloud database, extracts the stored content associated with each characteristic output text based on the characteristic output text generated by each spatial feature, and compiles each characteristic output text into an index text. Upon receiving user input, the query determination module 340 returns the associated requested data content based on the index text identified from the user input.
[0090] According to an optional embodiment, the method of confirming the feature difference based on the feature surface and multiple vertical surfaces to determine the selected surface includes: selecting a point to be used as a sound point to be confirmed based on the feature surface and multiple vertical surfaces, identifying the distance between the point to be sounded and the first feature point, and the distance between the point to be sounded and the second feature point, so as to calculate the first time feature value of the acquisition node corresponding to the first feature point, the second time feature value of the acquisition node corresponding to the second feature point, and confirming the feature difference to obtain a first selected surface determined by the first acquisition node and the second acquisition node; repeating the relevant calculation and determination steps to obtain a second selected surface determined based on the first acquisition node and the third acquisition node, and a third selected surface determined based on the second acquisition node and the third acquisition node.
[0091] According to an optional implementation manner, a time feature difference is calculated based on the first time feature difference and the second time feature value; the time difference corresponding to the two groups of associated moments are confirmed in pairs, with any one group of associated moments being the preceding feature and the other group of associated moments being the following feature, and the feature difference is confirmed based on the time difference between the preceding feature and the following feature; a plane having an equal feature difference and a time difference is selected or locked from multiple vertical planes, i.e., the first selected plane.
[0092] According to an optional embodiment, the intersection line formed by any two selected surfaces among the first selected surface, the second selected surface, and the third selected surface is used as the spatial feature of the relevant speech signal, and specifically the spatial position information of the formed intersection line is used as the spatial feature of the relevant speech signal, wherein the spatial position information includes the length of the intersection line, the angle formed by the intersection line and the vertical direction, the angle formed by the intersection line and the horizontal direction perpendicular to the vertical direction, and the distance between the intersection line and the specified point.
[0093] According to an optional embodiment, the confirming feature output text includes: receiving multiple groups of speech signals belonging to the same spatial feature, and confirming the audio features associated with the corresponding speech signals based on the audio signal waveforms associated with each of the different speech signals; and calculating the audio features of the speech signals using the following expression:
[0094] TT=JJ×C1+YY×C2
[0095] Among them, TT represents the audio characteristics of the speech signal; JJ represents the distance mean confirmed by averaging several groups of horizontal and vertical distances confirmed in the current audio signal waveform; YY represents the audio mean of the current audio signal confirmed by averaging several groups of audio frequencies; C1 and C2 are the first fixed coefficient factor and the second fixed coefficient factor respectively, the value of C1 is 0.5~0.7, and the value of C2 is 0.4~0.5.
[0096] According to an optional implementation, the following expression is used to calculate the comprehensive proportion of each text character to determine the proportion of the first group:
[0097] ZB1=G1÷G2
[0098] Among them, ZB1 represents the first group of proportion values calculated and confirmed, G1 represents the number of identical text characters recorded in the corresponding header of the recognition feature output text and the index text; G2 represents the total number of text characters recorded in the corresponding header of the recognition feature output text and the index text.
[0099] Use the following expression to calculate the comprehensive proportion of each text character and confirm the proportion of the second group:
[0100] ZB2=G1÷G3
[0101] Among them, ZB2 represents the second group of proportion values calculated and confirmed; G1 represents the number of identical text characters recorded in the recognition feature output text whose corresponding headers are the same as the index text; G3 represents the total number of index texts recorded in the recognition feature output text whose corresponding headers are the same as the index text.
[0102] The first group of proportion values and the second group of proportion values are averaged to determine the comprehensive proportion associated with the same text characters.
[0103] If there are multiple headers with the same text characters, the header associated with the largest comprehensive proportion is selected from the comprehensive proportions associated with each different header and recorded as the selected header, and the selected header is used as the feature header.
[0104] According to an optional implementation, multiple groups of audio features are averaged, the audio feature mean is locked, and the fluctuation range [YJ-Y2, YJ+Y2] is confirmed based on the preset value, where YJ is the locked audio feature mean; Y2 is a preset value, and the value of Y2 is 2 to 4; based on the fluctuation range, the fluctuation range with the largest total number of audio features is recorded as the determined range, and the speech signal associated with the abnormal feature is eliminated using the determined range.
[0105] It should be noted that due to Figure 3 The voice interaction optimization method performed by the voice interaction optimization device and Figure 1 The voice interaction optimization methods in the examples are roughly the same, so the description of the same parts is omitted.
[0106] Compared with the prior art, the present invention uses a combination of multiple groups of different sound signal acquisition nodes for processing. The spatial characteristics of the relevant voice signals are determined by determining different selected surfaces and finding the intersection lines formed by any two selected surfaces. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces positioning errors, and ensures that the source position of each voice signal can be accurately identified even in a complex multi-voice environment. According to the spatial characteristics of the voice signal, audio processing is performed on the voice signal in the same space, and the audio characteristics of the voice signal are determined by calculating the peak point, horizontal and vertical distance, and audio frequency of the audio signal waveform. Then, the continuously generated voice signal is averaged and analyzed for fluctuation intervals to eliminate voice signals with abnormal characteristics, effectively remove interference signals such as noise, and improve the quality and purity of the voice signal, making the subsequent text conversion more accurate and reliable. It can not only fully locate, but also effectively interact and remove noise.
[0107] Figure 4 is a schematic structural diagram of an electronic device according to an embodiment of the present invention.
[0108] like Figure 4 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.
[0109] The memory stores a computer executable program, typically a machine-readable code, which can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some of the steps in the method.
[0110] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).
[0111] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with an external device. The I / O interface may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0112] It should be understood that Figure 4 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute a computer-readable program stored in its memory to implement the method of the present invention or at least some of the steps of the method, it is considered an electronic device covered by the present invention.
[0113] Through the above description of the embodiments, it is easy for those skilled in the art to understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Figure 5 As shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.
[0114] The software product may utilize any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0115] The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. The data signal propagated may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0116] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0117] The computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the computer-readable medium implements the data interaction method of the present disclosure.
[0118] Those skilled in the art will appreciate that the modules described above can be distributed in the device according to the description of the embodiment, or can be modified accordingly to be used in one or more devices that are different from the embodiment. The modules of the above embodiment can be combined into one module or further divided into multiple submodules.
[0119] From the above description of the embodiments, those skilled in the art will readily appreciate that the exemplary embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored on a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes commands that cause a computing device (such as a personal computer, server, mobile terminal, or network device) to execute the methods according to the embodiments of the present invention.
[0120] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs.
[0121] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless the context dictates otherwise. The illustrated embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein.
[0122] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A voice interaction optimization method based on cross-positioning, characterized in that: include: Speech signals from at least three acquisition nodes are collected, and spatial position information associated with different speech signals is confirmed to obtain spatial features associated with each speech signal, specifically including: confirming the time difference corresponding to two groups of associated moments in a pairwise manner, determining a characteristic surface, and constructing multiple vertical surfaces perpendicular to the characteristic surface; confirming the characteristic difference based on the characteristic surface and the multiple vertical surfaces to determine a selected surface; and using the intersection line formed by any two selected surfaces as the spatial feature of the relevant speech signal; confirming the characteristic difference based on the characteristic surface and the multiple vertical surfaces to determine the selected surface includes: Based on the characteristic surface and multiple vertical surfaces, a point is selected to be used as a sounding point to be confirmed, and a distance between the sounding point to be confirmed and a first characteristic point, and a distance between the sounding point to be confirmed and a second characteristic point are identified to calculate a first time characteristic value of a collection node corresponding to the first characteristic point, a second time characteristic value of a collection node corresponding to the second characteristic point, and confirm a characteristic difference to obtain a first selected surface determined by the first collection node and the second collection node; Repeat the relevant calculation and determination steps to obtain a second selected surface determined based on the first acquisition node and the third acquisition node, and a third selected surface determined based on the second acquisition node and the third acquisition node. Based on the spatial features associated with each speech signal obtained, multiple groups of speech signals belonging to the same spatial feature are audio processed, and then the speech signals belonging to the same spatial feature are text-output, and the feature output text is confirmed; Based on a self-built cloud database, based on the feature output text generated by each spatial feature, the stored content associated with each feature output text is extracted, and each feature output text is compiled into an index text; the self-built cloud database includes each title, and each title corresponds to a storage area for storing the feature output text associated with the spatial feature; When user input is received, the associated requested data content is returned based on the index text recognized from the user input.
2. The cross-localization-based voice interaction optimization method according to claim 1, characterized in that: Further including: Calculating a time feature difference value based on the first time feature difference value and the second time feature value; Determine the time difference between two sets of associated moments in pairs, using any one set of associated moments as the preceding feature and the other set of associated moments as the following feature, and determine the feature difference based on the time difference between the preceding and following features; A plane having the same characteristic difference and time difference is selected or locked from multiple vertical planes, ie, the first selected plane.
3. The cross-positioning-based voice interaction optimization method according to claim 2, characterized in that: Further including: The intersection line formed by any two selected surfaces among the first selected surface, the second selected surface, and the third selected surface is used as the spatial feature of the relevant speech signal, and specifically the spatial position information of the formed intersection line is used as the spatial feature of the relevant speech signal, wherein, The spatial position information includes the length of the intersection line, the angle formed by the intersection line and the vertical direction, the angle formed by the intersection line and the horizontal direction perpendicular to the vertical direction, and the distance between the intersection line and the designated point.
4. The cross-localization-based voice interaction optimization method according to claim 1, characterized in that: The confirmation feature output text includes: receiving multiple groups of speech signals belonging to the same spatial feature, and determining audio features associated with the corresponding speech signals based on audio signal waveforms associated with the different speech signals; Use the following expression to calculate the audio features of the speech signal: TT=JJ×C1+YY×C2 Among them, TT represents the audio characteristics of the speech signal; JJ represents the distance mean confirmed by averaging several groups of horizontal and vertical distances confirmed in the current audio signal waveform; YY represents the audio mean of the current audio signal confirmed by averaging several groups of audio frequencies; C1 and C2 are the first fixed coefficient factor and the second fixed coefficient factor respectively, the value of C1 is 0.5~0.7, and the value of C2 is 0.4~0.
5.
5. The cross-localization-based voice interaction optimization method according to claim 1, characterized in that: Further including: Use the following expression to calculate the comprehensive proportion of each text character and confirm the proportion of the first group: ZB1=G1÷G2 Where ZB1 represents the first group of percentage values calculated and confirmed, G1 represents the number of identical text characters recorded in the corresponding header of the recognition feature output text and the index text; G2 represents the total number of text characters in the corresponding header of the recognition feature output text and the index text. Use the following expression to calculate the comprehensive proportion of each text character and confirm the proportion of the second group: ZB2=G1÷G3 Wherein, ZB2 represents the second group of percentage values calculated and confirmed; G1 represents the number of identical text characters recorded in the corresponding header of the recognition feature output text and the index text; G3 represents the total number of index texts recorded in the corresponding header of the recognition feature output text and the index text; The first group of proportion values and the second group of proportion values are averaged to determine the comprehensive proportion associated with the same text characters; If there are multiple headers with the same text characters, the header associated with the largest comprehensive proportion is selected from the comprehensive proportions associated with each different header and recorded as the selected header, and the selected header is used as the feature header.
6. The cross-positioning-based voice interaction optimization method according to claim 1, characterized in that: Further including: Performing averaging on multiple sets of audio features, locking the audio feature mean, and determining the fluctuation range [YJ-Y2, YJ+Y2] based on a preset value, where YJ is the locked audio feature mean; Y2 is the preset value, and the value of Y2 ranges from 2 to 4; Based on the fluctuation interval, the fluctuation interval in which the total number of audio features is the largest is recorded as a determination interval, and the speech signal associated with the abnormal feature is eliminated using the determination interval.
7. A voice interaction optimization device based on cross-positioning, characterized in that: It executes the cross-positioning-based voice interaction optimization method according to any one of claims 1 to 6, and the voice interaction optimization device includes: The acquisition and processing module is configured to acquire speech signals from at least three acquisition nodes and determine the spatial position information associated with different speech signals to obtain spatial features associated with each speech signal. Specifically, the module determines the time difference between two groups of associated moments in a pairwise manner, determines a characteristic surface, and constructs multiple vertical surfaces perpendicular to the characteristic surface; determines the characteristic difference based on the characteristic surface and the multiple vertical surfaces to determine a selected surface; and uses the intersection line formed by any two selected surfaces as the spatial feature of the relevant speech signal. A confirmation module, based on the spatial features associated with each speech signal, performs audio processing on multiple groups of speech signals belonging to the same spatial feature, then outputs text for the speech signals belonging to the same spatial feature, and confirms the feature output text; The extraction and processing module, based on the self-built cloud database, extracts the storage content associated with each feature output text according to the feature output text generated by each spatial feature, and compiles each feature output text into an index text; The query determination module, when receiving user input, returns associated requested data content based on index text recognized from the user input.
8. The cross-positioning-based voice interaction optimization device according to claim 7, characterized in that: include: Based on the characteristic surface and multiple vertical surfaces, a point is selected to be used as a sounding point to be confirmed, and a distance between the sounding point to be confirmed and a first characteristic point, and a distance between the sounding point to be confirmed and a second characteristic point are identified to calculate a first time characteristic value of a collection node corresponding to the first characteristic point, a second time characteristic value of a collection node corresponding to the second characteristic point, and confirm a characteristic difference to obtain a first selected surface determined by the first collection node and the second collection node; The related calculation and determination steps are repeatedly performed to obtain a second selected surface determined based on the first acquisition node and the third acquisition node, and a third selected surface determined based on the second acquisition node and the third acquisition node.
9. The cross-positioning-based voice interaction optimization device according to claim 8, characterized in that: include: Calculating a time feature difference value based on the first time feature difference value and the second time feature value; Determine the time difference between two sets of associated moments in pairs, using any one set of associated moments as the preceding feature and the other set of associated moments as the following feature, and determine the feature difference based on the time difference between the preceding and following features; A plane having the same characteristic difference and time difference is selected or locked from multiple vertical planes, ie, the first selected plane.
Citation Information
Patent Citations
Equipment control method and device, storage medium and electronic device
CN113450798A
Display method and device, voice equipment and storage medium
CN115705849A