Voice interaction optimization method and device based on cross positioning
Through the combined cross-positioning technology of multiple sets of sound signal acquisition nodes, the problem of inaccurate voice signal positioning in the voice interaction system is solved, high-precision voice signal recognition and effective interaction in complex environments are achieved, and the speech signal quality and text conversion accuracy are improved.
Patent Information
- Application Number
- CN202510764739.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing voice interaction technology is difficult to accurately identify the source direction and location of the voice signal in complex scenarios where multiple people make vocals at the same time, resulting in the inability to effectively distinguish the voice signals of different people, affecting the accuracy and efficiency of interaction.
Using multiple sets of different sound signal acquisition nodes, the spatial characteristics of the speech signal are determined by determining the selected surface and finding the intersection lines formed by any two selected surfaces, and the audio processing is performed by calculating the peak point, lateral vertical distance and audio frequency of the audio signal, and abnormal signals are eliminated to improve the quality and purity of the speech signal.
Accurately identifying the source location of each voice signal in a complex multi-speech environment improves the reliability and accuracy of positioning, ensures the quality and purity of the voice signal, and makes subsequent text conversions more accurate and reliable, and can effectively remove noise.
Smart Images

Figure CN120279900A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voice interaction, and provides a voice interaction optimization method and device based on cross-positioning. Background Art
[0002] With the rapid development of artificial intelligence technology, voice interaction, as a natural and convenient way of human-computer interaction, is being widely used in many fields such as smart home, smart customer service, in-vehicle systems, smart wearable devices, etc. However, in actual application scenarios, voice interaction faces many challenges, especially in complex cross-scenario environments, the existing voice interaction technology has certain limitations.
[0003] Traditional voice interaction systems often lack the ability to accurately identify the spatial location of voice signals. Most systems can simply receive voice input, but cannot clearly identify the source direction and specific location of the voice signal. In complex scenarios where multiple people speak at the same time, such as conference rooms and restaurants, voice signals from different locations are intertwined. Existing technologies are difficult to accurately distinguish the spatial characteristics of each voice signal, resulting in the inability to effectively process the voices of different people in a targeted manner, affecting the accuracy and efficiency of the interaction.
[0004] Therefore, it is necessary to provide a voice interaction optimization method based on cross-localization to solve the above problems. Summary of the invention
[0005] The present invention provides a voice interaction optimization method and device based on cross-positioning to solve the problem in the prior art that the source direction and specific location of the voice signal cannot be clearly identified, especially in complex scenarios where multiple people speak at the same time (such as conference rooms, restaurants, etc. where voice signals from different locations are intertwined), it is difficult to accurately distinguish the spatial characteristics of each voice signal, resulting in the inability to effectively perform targeted processing on the voices of different people, affecting the accuracy and efficiency of the interaction. The technical problem to be solved by the present invention is achieved through the following technical solution.
[0006] In a first aspect of the present invention, an optimization method for voice interaction based on cross - positioning is proposed. The method includes: collecting voice signals of at least three collection nodes, and confirming the spatial position information associated with different voice signals to obtain the spatial features associated with each voice signal. Specifically, it includes: confirming the time difference corresponding to two sets of associated moments in pairs, determining a feature plane, and constructing a plurality of vertical planes perpendicular to the feature plane; based on the feature plane and the plurality of vertical planes, confirming a feature difference to determine a selected plane; using the intersection line formed by any two selected planes as the spatial feature of the relevant voice signal; based on the obtained spatial features associated with each voice signal, performing audio processing on multiple groups of voice signals belonging to the same spatial feature, and then performing text output on the voice signals belonging to the same spatial feature to confirm a feature output text; based on a self - built cloud database, extracting the stored content associated with each feature output text according to the feature output text generated by each spatial feature, and compiling each feature output text into an index text; when receiving a user input, based on the index text recognized from the user input, returning the associated requested data content.
[0007] In a second aspect of the present invention, an optimization device for voice interaction based on cross - positioning is proposed. It executes the optimization method for voice interaction based on cross - positioning described in the first aspect of the present invention. The voice interaction optimization device includes: a collection and processing module, configured to collect voice signals of at least three collection nodes, and confirm the spatial position information associated with different voice signals to obtain the spatial features associated with each voice signal. Specifically, it includes: confirming the time difference corresponding to two sets of associated moments in pairs, determining a feature plane, and constructing a plurality of vertical planes perpendicular to the feature plane; based on the feature plane and the plurality of vertical planes, confirming a feature difference to determine a selected plane; using the intersection line formed by any two selected planes as the spatial feature of the relevant voice signal; a confirmation module, based on the obtained spatial features associated with each voice signal, performing audio processing on multiple groups of voice signals belonging to the same spatial feature, and then performing text output on the voice signals belonging to the same spatial feature to confirm a feature output text; an extraction and processing module, based on a self - built cloud database, extracting the stored content associated with each feature output text according to the feature output text generated by each spatial feature, and compiling each feature output text into an index text; a query and determination module, when receiving a user input, based on the index text recognized from the user input, returning the associated requested data content.
[0008] In a third aspect of the present invention, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the optimization method for voice interaction based on cross - positioning described in the first aspect of the present invention.
[0009] A fourth aspect of the present invention provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for optimizing voice interaction based on cross-positioning described in the first aspect of the present invention is implemented.
[0010] The embodiments of the present invention include the following advantages: Compared with the prior art, the present invention processes by using multiple groups of different sound signal acquisition node combinations, determines the spatial characteristics of relevant voice signals by determining different selected planes and finding the intersection lines formed by any two selected planes. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces the positioning error, and ensures that the source positions of each voice signal can be accurately identified even in a complex multi-voice environment. Audio processing is performed on the voice signals in the same space according to the spatial characteristics of the voice signals. By calculating features such as the peak points, horizontal vertical distances, and audio frequencies of the audio signal waveforms, the audio characteristics of the voice signals are determined. Then, mean processing and fluctuation range analysis are performed on the continuously generated voice signals to eliminate the voice signals with abnormal characteristics, effectively removing interference signals such as noise, improving the quality and purity of the voice signals, making the subsequent text conversion more accurate and reliable, and not only being able to fully position but also effectively interact and perform noise removal. Description of the Drawings
[0011] Figure 1 is a flowchart of the steps of an example of the method for optimizing voice interaction based on cross-positioning of the present invention; Figure 2 is a schematic diagram of the principle from another angle of the method for optimizing voice interaction based on cross-positioning of the present invention; Figure 3 is a block diagram of the structure of the device for optimizing voice interaction based on cross-positioning of the present invention; Figure 4 is a schematic diagram of the structure of an embodiment of an electronic device according to the present invention; Figure 5 is a schematic diagram of the structure of an embodiment of a computer-readable medium according to the present invention. Detailed Embodiments
[0012] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0013] In traditional voice interaction systems, there is often a lack of the ability to accurately identify the spatial position of voice signals. They can only simply receive voice input and cannot determine the source direction and specific location of the voice signals. In complex scenarios where multiple people speak simultaneously, such as in conference rooms, restaurants, etc., voice signals from different positions are intertwined, and existing technologies are difficult to accurately distinguish the spatial characteristics of each voice signal, resulting in the inability to effectively process the voices of different people targeted, affecting the accuracy and efficiency of interaction.
[0014] In view of the above problems, the present invention proposes an optimized method for voice interaction based on cross-location. This method uses a combination of multiple groups of different sound signal acquisition nodes for processing. By determining different selected planes and finding the intersection lines formed by any two selected planes, the spatial characteristics of the relevant voice signals are determined. This multi-group combination cross-location method further improves the reliability and accuracy of location, effectively reduces the location error, and ensures that the source location of each voice signal can be accurately identified even in a complex multi-voice environment. According to the spatial characteristics of the voice signals, audio processing is performed on the voice signals in the same space. By calculating features such as the peak points, horizontal vertical distances, and audio frequencies of the audio signal waveforms, the audio characteristics of the voice signals are determined. Then, mean processing and fluctuation range analysis are performed on the continuously generated voice signals to eliminate voice signals with abnormal characteristics, effectively removing interference signals such as noise, improving the quality and purity of the voice signals, making the subsequent text conversion more accurate and reliable, and not only being able to fully locate but also effectively interact and perform noise removal.
[0015] Embodiment 1 The following will refer to Figure 1 、 Figure 2 to describe the content of the present invention in detail.
[0016] Figure 1 is a step flowchart of an example of the optimized method for voice interaction based on cross-location of the present invention.
[0017] As Figure 1 shown, in step S101, voice signals of at least three acquisition nodes are collected, and the spatial position information associated with different voice signals is confirmed to obtain the spatial characteristics associated with each voice signal.
[0018] Specifically, collect the voice signals of at least three collection nodes. For example, there are three collection nodes, specifically including the first collection node (such as a voice interaction device, specifically including a collector, a smart speaker, a car speaker, a wearable device), the second collection node and the third collection node arranged on the top and / or side of the voice interaction device. As long as the second collection node and the third collection node are arranged asymmetrically, and the three groups of sound signal collection points of the first collection node, the second collection node, and the third collection node are non-symmetrically arranged, thus forming a staggered or cross-shaped pattern, that is, arranged in a staggered or cross-shaped pattern, which can more effectively lock the features and facilitate the subsequent specific confirmation of the spatial position.
[0019] According to the collected voice signals, confirm the spatial position information associated with different voice signals, specifically including the following steps: Step S201: Confirm the associated moments of the voice signals received from different collection nodes, confirm the time differences corresponding to the two groups of associated moments in pairs, use any group of associated moments as the pre-feature, use the other group of associated moments as the post-feature, and use the two collection nodes associated with the time difference between the pre-feature and the post-feature as the first feature point and the second feature point, and determine the feature plane based on the first feature point and the second feature point, and construct a plurality of vertical planes perpendicular to the feature plane.
[0020] Specifically, for a voice signal, there are three collected associated moments t1, t2, and t3, and there are time differences |t1 - t2|, |t2 - t3|, |t1 - t3| between each moment.
[0021] Confirm the associated moments when the collection nodes of different sound signals receive the corresponding voice signals, and mark them as Si, where i represents different sound signal collection nodes, that is, the i-th collection node, and i is a positive integer. In the present invention, specifically 1, 2, 3. In pairs, confirm the time differences of the associated moments corresponding to the two different sound signal collection nodes, use any group of associated moments as the pre-feature T1, use the other group of associated moments as the post-feature T2, and use the time difference = T1 - T2 to confirm the time difference associated with the corresponding sound signal collection node.
[0022] Then, use the two collection nodes of the two different voice signals associated with the above time difference, specifically the first collection node and the second collection node, as the first feature point and the second feature point.
[0023] For example, Figure 2The midpoint A, point B, and point C are respectively used as the first feature point, the second feature point, and the third feature point. Based on the first feature point and the second feature point, a feature plane is determined. Specifically, points A and B are connected to form a line segment AB passing through points A and B, and the plane corresponding to the line segment AB is used as the feature plane, and multiple vertical planes perpendicular to the feature plane are constructed. Among them, point A, point B, and point C respectively correspond to the first acquisition point, the second acquisition point, and the third acquisition point.
[0024] Step S202: Based on the feature plane and multiple vertical planes, select points to be used as the to-be-confirmed sound emission points, identify the distances between the to-be-emitted sound points and the first feature point, and the distances between the to-be-emitted sound points and the second feature point, so as to calculate the first time feature value of the acquisition node corresponding to the first feature point and the second time feature value of the acquisition node corresponding to the second feature point, further confirm the feature difference value, and determine the selected plane.
[0025] Randomly select a set of points from different vertical planes, denoted as the to-be-confirmed sound emission points, identify the straight-line distances L1 and L2 between the to-be-confirmed sound emission points and the first feature point and the second feature point respectively, and then use L1÷c = t1 and L2÷c = t2 to confirm the first time feature value t1 and the second time feature value t2 associated with the first acquisition node and the second acquisition node of two different sound signals. Among them, c is the preset propagation speed of sound in the air, for example, 343 m / s.
[0026] Based on the confirmed time feature values t1 and t2, place the time feature value associated with the prefeature T1 in the front and the time feature value associated with the postfeature T2 in the back, perform difference processing, and confirm the feature difference value. Select or lock the plane from multiple vertical planes where the feature difference value is equal to the time difference value, that is, the first selected plane (such as Figure 2 the "selected plane" shown in
[0027] Step S203: Repeat the relevant steps S201 and S202 (i.e., the calculation and determination steps) to obtain the second selected plane determined based on the first acquisition node and the third acquisition node, and the third selected plane determined based on the second acquisition node and the third acquisition node.
[0028] It should be noted that there are three acquisition nodes on the corresponding acquisition device, and there is a set of planes where each two acquisition nodes are located, that is, the corresponding feature planes. Several vertical planes can be confirmed on the surface of the feature plane. When the time features associated with the corresponding points in the vertical plane are consistent with the time features associated with the third acquisition node, that is, the position of the selected plane, the sound emission point is located in the corresponding selected plane. The above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.
[0029] Step S204: Use the intersection line formed by any two of the first selected plane, the second selected plane, and the third selected plane as the spatial feature of the relevant voice signal.
[0030] Specifically, identify the intersection line generated by any two selected planes in space (since the selected planes cannot be parallel, there must be an intersection line), and use the spatial position information of the generated intersection line as the spatial feature of the relevant voice signal. Specifically, the spatial position information includes the length of the intersection line, the angle formed by the intersection line and the vertical direction (or the horizontal direction perpendicular to the vertical direction), the distance between the intersection line and a specified point (such as the first acquisition node, the second acquisition node, the third acquisition node, or other points, etc.), and so on.
[0031] By confirming the spatial position information associated with different voice signals, the spatial features associated with each voice signal are obtained. Specifically, by determining the intersection line generated by the selected plane as the spatial feature associated with the relevant voice signal, this multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces the positioning error, and ensures that the source position of each voice signal can be accurately identified in a complex multi-voice environment.
[0032] It should be noted that the above is only an optional example for illustration and should not be construed as a limitation to the present invention.
[0033] Next, in step S102, based on the spatial features associated with each voice signal obtained, perform audio processing on multiple groups of voice signals belonging to the same spatial feature, and then perform text output on the voice signals belonging to the same spatial feature to confirm the feature output text.
[0034] In a specific embodiment, for voice signal a and voice signal b belonging to the same spatial feature (such as intersection line x), perform audio processing on voice signal a and voice signal b, initially eliminate the relevant voice signals with inconsistent audio features, and then perform feature text output on voice signal a and voice signal b through audio processing to confirm the feature output text.
[0035] In the case where multiple voice signals correspond to the same spatial feature, the audio features associated with each different voice signal are different. By confirming the audio features, different voice signals are distinguished.
[0036] For confirming the feature output text, it specifically includes the following steps: Step S301: Receive multiple groups of voice signals belonging to the same spatial feature, and based on the audio signal waveforms associated with the different voice signals respectively, confirm the audio features associated with the corresponding voice signals.
[0037] For example, voice signals a and b belonging to the same spatial feature (such as cross line x) are received, and based on the audio signal waveforms associated with voice signal a and voice signal b respectively.
[0038] Based on the audio signal waveforms associated with voice signal a and voice signal b respectively, for a single set of voice signals (such as voice signal a or voice signal b), from the current audio signal waveform, the peak points existing in the current audio signal waveform are identified, the horizontal vertical distances between adjacent peak points are recorded and represented by Fk, where k represents different adjacent peak points. Then, the audio frequencies associated with each set of peak points are identified and represented by Pk. The several sets of horizontal vertical distances Fk identified in the current audio signal waveform are averaged to obtain the distance average JJ. Then, the several sets of audio frequencies Pk are averaged to obtain the audio average YY of the current audio signal.
[0039] The following expression is used to calculate the audio feature of the voice signal: TT = JJ × C1 + YY × C2 Where TT represents the audio feature of the voice signal; JJ represents the distance average obtained by averaging several sets of horizontal vertical distances identified in the current audio signal waveform; YY represents the audio average of the current audio signal obtained by averaging several sets of identified audio frequencies; C1 and C2 are the first fixed coefficient factor and the second fixed coefficient factor respectively. The value of C1 ranges from 0.5 to 0.7 and is 0.574 in this example. The value of C2 ranges from 0.4 to 0.5 and is 0.426 in this example.
[0040] Step S302: Perform audio processing on multiple sets of voice signals continuously generated for the same spatial feature, identify the audio features associated with different voice signals, and reprocess several sets of audio features TT.
[0041] The so-called continuity means that the time difference between different voice signals is less than 1 second, and this time difference represents the continuous generation of the corresponding voice signals.
[0042] Several groups of audio features TT are reprocessed, specifically including averaging multiple groups of audio features, locking the audio feature mean, confirming the fluctuation range [YJ-Y2, YJ+Y2] based on the preset value Y2, where YJ is the locked audio feature mean, Y2 is the preset value, and the value of Y2 is 2 to 4, for example, prepared in advance by relevant operators based on experience, keeping the range of the above-mentioned fluctuation range unchanged, making the endpoint values of the fluctuation range change synchronously, and recording the audio features TT included in the fluctuation range corresponding to each different numerical range, recording the fluctuation range with the largest total number of audio features as the determined interval, and recording the audio features TT that do not belong to the determined interval as abnormal features, eliminating the voice signals associated with the abnormal features, and re-sorting the multiple groups of continuously generated voice signals currently detected, and confirming the voice signal sorting sequence.
[0043] Next, the voice signal sorting sequence is subjected to analog-to-digital processing to lock the digital signal associated with the corresponding voice signal, and then the associated characters are locked based on the digital signal. The characters that are locked in sequence are kept unchanged according to the original sorting method to generate a feature output text associated with the corresponding spatial feature. For example: "What's the temperature today?", "Today" corresponds to a digital signal, "Day" corresponds to a digital signal, and so on. The corresponding characters are gradually confirmed to determine the corresponding output text, thereby locking the feature output text: "What's the temperature today?"
[0044] Specifically, the corresponding voice content is generated into text content, and the relevant conversion can be performed automatically. The abnormal voice signal associated with the corresponding spatial feature is generally noise, and its signal feature is greatly different from other signal features. Specifically, these abnormal audios such as noise are removed.
[0045] It should be noted that the above is only described as an optional example and should not be understood as a limitation to the present invention.
[0046] Next, in step S103, based on the self-built cloud database, according to the feature output texts generated by each spatial feature, the storage content associated with each feature output text is extracted, and each feature output text is compiled into an index text.
[0047] Specifically, the self-built cloud database includes various titles (ie, headers), and each title corresponds to a storage area for storing feature output texts associated with spatial features.
[0048] It should be noted that, in this example, the header can be understood as the title of a storage area, and for each title, there is a corresponding stored text output content.
[0049] S41. Label the different feature output texts associated with different spatial features (i.e., the spatial features associated with each voice signal obtained in step S102) as index texts, and extract the headers of other output data (i.e., the corresponding index texts) from the cloud database. The headers are all preset by relevant personnel in advance. Compare and verify the index texts with different headers to lock the feature headers.
[0050] Specifically, it includes the following steps.
[0051] S411: Identify the text characters in the feature output text where the corresponding header is the same as the index text, and record the number G1 of the same text characters. Then record the total number G2 of text characters in the corresponding header and the total number G3 of the index text, and confirm the comprehensive proportion of the same text characters.
[0052] Preferably, use the following expression to calculate the comprehensive proportion of each text character and confirm the first group proportion: ZB1 = G1÷G2 Among them, ZB1 represents the calculated and confirmed first group proportion value, G1 represents the number of the same text characters identified in the feature output text where the corresponding header is the same as the index text; G2 represents the total number of text characters in the corresponding header identified in the feature output text where the corresponding header is the same as the index text.
[0053] Use the following expression to calculate the comprehensive proportion of each text character and confirm the second group proportion: ZB2 = G1÷G3 Among them, ZB2 represents the calculated and confirmed second group proportion value; G1 represents the number of the same text characters identified in the feature output text where the corresponding header is the same as the index text; G3 represents the total number of the index text identified in the feature output text where the corresponding header is the same as the index text.
[0054] Specifically, perform mean processing on the first group proportion value ZB1 and the second group proportion value ZB2 to lock the comprehensive proportion associated with the same text characters.
[0055] S412: If there is no header with the same text characters, directly generate an error signal for display.
[0056] If there is a header with a group of the same text characters, label this header as the feature header.
[0057] If there are multiple headers with the same text characters, select the header associated with the largest comprehensive proportion from the comprehensive proportions associated with each different header and record it as the selected header, and use the selected header as the feature header.
[0058] For example, assume the determined characteristic output text is: "What's the temperature today?". In its cloud database, there are headers corresponding to the output data: the headers can be "Today's temperature", "Yesterday's temperature", "The day before yesterday's temperature", "City temperature", etc., or other headers. Step S42: Confirm the index path of the characteristic header in the characteristic output text, and perform voice output on the stored content associated with the determined index path to complete real-time interaction.
[0059] For example, assume the determined characteristic output text is: "What's the temperature today?"; In the cloud database, there are headers stored for the corresponding output data, specifically including the following headers: "Today's temperature", "Yesterday's temperature", "The day before yesterday's temperature", "City temperature", etc., or other headers.
[0060] It should be noted that in this example, the header can be understood as the title of a storage area. For this title, there is corresponding stored output content. If there is only a single header (i.e., there are identical text characters), the content can be directly displayed; if there are multiple headers, then select "Today's temperature" as the characteristic header and perform relevant content display output.
[0061] Based on the header and the characteristic output text, it can be determined that "What's the temperature today?" has the highest overlap with the header "Today's temperature", that is, the determined proportion of the first group is 4 / 7, the proportion of the second group is 1, and the confirmed comprehensive proportion is 5.5 / 7.
[0062] If there is only a single header (i.e., there are identical text characters), the content can be directly displayed.
[0063] If there are multiple headers, then select one of the headers (e.g., "Today's temperature") as the characteristic header and perform relevant content display output.
[0064] It should be noted that the above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.
[0065] Next, in step S104, when receiving user input, based on the index text recognized from the user input, return the associated requested data content.
[0066] For example, if the user inputs "What's the temperature today?", recognize the index text (e.g., "Today's temperature") from the user input and return the associated requested data content: such as information like 25 degrees.
[0067] It should be noted that the above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.
[0068] Compared with the prior art, the present invention processes by combining multiple groups of different voice signal acquisition nodes. By determining different selected planes and finding the intersection lines formed by any two selected planes, the spatial characteristics of relevant voice signals are determined. This multi-group combined cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces the positioning error, and ensures that the source positions of each voice signal can be accurately identified even in a complex multi-voice environment. According to the spatial characteristics of the voice signals, audio processing is performed on the voice signals in the same space. By calculating features such as the peak points, horizontal vertical distances, and audio frequencies of the audio signal waveforms, the audio characteristics of the voice signals are determined. Then, mean processing and fluctuation range analysis are performed on the continuously generated voice signals to eliminate the voice signals with abnormal characteristics, effectively removing interference signals such as noise, improving the quality and purity of the voice signals, making the subsequent text conversion more accurate and reliable, and not only enabling full positioning but also effective interaction and noise removal.
[0069] Embodiment 2 The following is an embodiment of the device of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment of the present invention, please refer to the method embodiment of the present invention.
[0070] Figure 3 It is a schematic structural diagram of an example of a voice interaction optimization device based on cross-positioning according to the present invention. The following will be described with reference to Figure 3 , the voice interaction optimization device. The voice interaction optimization device is used to execute the method for voice interaction optimization described in the first aspect of the present invention.
[0071] As Figure 3 shown, the voice interaction optimization device 300 includes an acquisition and processing module 310, a confirmation module 320, an extraction and processing module 330, and a query and determination module 340.
[0072] In a specific embodiment, the acquisition and processing module 310 is configured to acquire voice signals of at least three acquisition nodes, confirm the spatial position information associated with different voice signals, so as to obtain the spatial features associated with each voice signal, specifically including: confirming the time difference corresponding to two sets of associated moments in a pairwise manner, determining a feature plane, and constructing a plurality of vertical planes perpendicular to the feature plane; based on the feature plane and the plurality of vertical planes, confirming a feature difference to determine a selected plane; using the intersection line formed by any two selected planes as the spatial feature of the relevant voice signal. The confirmation module 320 performs audio processing on multiple groups of voice signals belonging to the same spatial feature based on the obtained spatial features associated with each voice signal, and then outputs text for the voice signals belonging to the same spatial feature to confirm the feature output text. The extraction and processing module 330 extracts the stored content associated with each feature output text based on the self-built cloud database according to the feature output text generated by each spatial feature, and compiles each feature output text into an index text. When receiving a user input, the query determination module 340 returns the associated requested data content based on the index text identified from the user input.
[0073] According to an alternative embodiment, the step of confirming a feature difference based on the feature plane and the plurality of vertical planes to determine a selected plane includes: based on the feature plane and the plurality of vertical planes, selecting a point position to be used as the point to be confirmed for sound emission, identifying the distance between the point to be sounded and a first feature point, and the distance between the point to be sounded and a second feature point, so as to calculate a first time feature value of the acquisition node corresponding to the first feature point and a second time feature value of the acquisition node corresponding to the second feature point, and confirming a feature difference to obtain a first selected plane determined by the first acquisition node and the second acquisition node; repeating the relevant calculation and determination steps to obtain a second selected plane determined by the first acquisition node and the third acquisition node, and a third selected plane determined by the second acquisition node and the third acquisition node.
[0074] According to an alternative embodiment, a time feature difference is calculated based on the first time feature difference and the second time feature value; the time difference corresponding to two sets of associated moments is confirmed in a pairwise manner, with any set of associated moments as the pre-feature and the other set of associated moments as the post-feature, and the feature difference is confirmed according to the time difference between the pre-feature and the post-feature; a plane with an equal feature difference and time difference is selected or locked from the plurality of vertical planes, i.e., the first selected plane.
[0075] According to an alternative embodiment, the intersection line formed by any two of the first selected surface, the second selected surface, and the third selected surface is used as the spatial feature of the relevant voice signal. Specifically, the spatial position information of the formed intersection line is used as the spatial feature of the relevant voice signal, where the spatial position information includes the length of the intersection line, the angle formed by the intersection line and the vertical direction, the angle formed by the intersection line and the horizontal direction perpendicular to the vertical direction, and the distance between the intersection line and the specified point.
[0076] According to an alternative embodiment, the confirmation feature output text includes: receiving multiple groups of voice signals belonging to the same spatial feature, and based on the audio signal waveforms associated with different voice signals respectively, confirming the audio features associated with the corresponding voice signals; using the following expression to calculate the audio features of the voice signals: TT = JJ × C1 + YY × C2 where TT represents the audio feature of the voice signal; JJ represents the mean value of the confirmed several groups of horizontal vertical distances in the current audio signal waveform; YY represents the audio mean value of the current audio signal obtained by averaging the confirmed several groups of audio frequencies; C1 and C2 are the first fixed coefficient factor and the second fixed coefficient factor respectively, and the value of C1 ranges from 0.5 to 0.7, and the value of C2 ranges from 0.4 to 0.5.
[0077] According to an alternative embodiment, the comprehensive proportion of each text character is calculated using the following expression to confirm the first group proportion: ZB1 = G1 ÷ G2 where ZB1 represents the calculated and confirmed first group proportion value; G1 represents the number of the same text characters recorded for the text characters in the recognition feature output text where the corresponding header is the same as the index text; G2 represents the total number of text characters in the corresponding header in the recognition feature output text where the corresponding header is the same as the index text.
[0078] The comprehensive proportion of each text character is calculated using the following expression to confirm the second group proportion: ZB2 = G1 ÷ G3 where ZB2 represents the calculated and confirmed second group proportion value; G1 represents the number of the same text characters recorded for the text characters in the recognition feature output text where the corresponding header is the same as the index text; G3 represents the total number of the index text in the recognition feature output text where the corresponding header is the same as the index text.
[0079] The mean value of the first group proportion value and the second group proportion value is processed to determine the comprehensive proportion associated with the same text characters.
[0080] If there are multiple headers with the same text characters, then from the comprehensive ratios associated with each different header, select the header associated with the maximum comprehensive ratio as the selected header, and use the selected header as the characteristic header.
[0081] According to an alternative embodiment, perform mean processing on multiple groups of audio features, lock the mean value of the audio features, and confirm the fluctuation range [YJ - Y2, YJ + Y2] based on a preset value, where YJ is the locked mean value of the audio features; Y2 is the preset value, and the value of Y2 ranges from 2 to 4; based on the fluctuation range, record the fluctuation range with the largest number of audio features as the determined range, and use the determined range to eliminate the speech signals associated with abnormal features.
[0082] It should be noted that since Figure 3 the voice interaction optimization method executed by the voice interaction optimization device of Figure 1 is substantially the same as the voice interaction optimization method in the example of
[0083] Compared with the prior art, the present invention uses a combination of multiple groups of different sound signal acquisition nodes for processing, determines the spatial characteristics of relevant speech signals by determining different selected surfaces and finding the intersection lines formed by any two selected surfaces. This multi-group combination cross-positioning method further improves the reliability and accuracy of positioning, effectively reduces the positioning error, and ensures that the source positions of each speech signal can be accurately identified even in a complex multi-voice environment. Perform audio processing on the speech signals in the same space according to the spatial characteristics of the speech signals, and determine the audio characteristics of the speech signals by calculating features such as the peak points, horizontal vertical distances, and audio frequencies of the audio signal waveforms. Then perform mean processing and fluctuation range analysis on the continuously generated speech signals, and eliminate the speech signals with abnormal features, effectively removing interference signals such as noise, improving the quality and purity of the speech signals, making the subsequent text conversion more accurate and reliable, and not only being able to fully position but also effectively interact and perform noise removal.
[0084] Figure 4 is a schematic structural diagram of an embodiment of an electronic device according to the present invention.
[0085] As Figure 4 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work collaboratively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity and can also be the sum of multiple physical devices.
[0086] The memory stores computer-executable programs, typically machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to execute the method of the present invention or at least some of the steps in the method.
[0087] The memory includes volatile memory, such as random access storage units (RAM) and / or cache storage units, and may also include non-volatile memory, such as read-only storage units (ROM).
[0088] Optionally, in this embodiment, the electronic device further includes an I / O interface for data exchange between the electronic device and external devices. The I / O interface can represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.
[0089] It should be understood that Figure 4 The electronic device shown is merely an example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above example. For example, some electronic devices also include a display unit such as a display screen, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least some of the steps of the method, it can be considered as an electronic device covered by the present invention.
[0090] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, as Figure 5 shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.
[0091] The software product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0092] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0093] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partially on the user's device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0094] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by a device, the computer-readable medium implements the data interaction method of the present disclosure.
[0095] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are only different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0096] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.
[0097] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0098] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, like symbols typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.
[0099] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An optimized method for voice interaction based on cross-positioning, characterized in that, Including: Collect the voice signals of at least three collection nodes, confirm the spatial position information associated with different voice signals, so as to obtain the spatial features associated with each voice signal, specifically including: confirm the time difference corresponding to two groups of associated moments in a pairwise manner, determine the feature plane, and construct multiple vertical planes perpendicular to the feature plane; based on the feature plane and multiple vertical planes, confirm the feature difference to determine the selected plane; use the intersection line formed by any two selected planes as the spatial feature of the relevant voice signal; Based on the spatial features associated with each obtained voice signal, perform audio processing on multiple groups of voice signals belonging to the same spatial feature, and then perform text output on the voice signals belonging to the same spatial feature to confirm the feature output text; Based on the self-built cloud database, extract the stored content associated with each feature output text according to the feature output text generated by each spatial feature, and compile each feature output text into an index text; When receiving a user input, based on the index text recognized from the user input, return the associated requested data content.
2. The voice interaction optimization method based on cross positioning according to claim 1, wherein the step of confirming the feature difference based on the feature plane and multiple vertical planes to determine the selected plane includes: Based on the feature plane and multiple vertical planes, select points to be used as the points to be confirmed for sound emission, identify the distance between the point to be sounded and the first feature point, and the distance between the point to be sounded and the second feature point, so as to calculate the first time feature value of the collection node corresponding to the first feature point and the second time feature value of the collection node corresponding to the second feature point, and confirm the feature difference to obtain the first selected plane determined by the first collection node and the second collection node; Repeatedly execute the relevant calculation and determination steps to obtain the second selected plane determined by the first collection node and the third collection node, and the third selected plane determined by the second collection node and the third collection node.
3. The voice interaction optimization method based on cross positioning according to claim 2, characterized in that Further including: Calculate the time feature difference according to the first time feature difference and the second time feature value; Confirm the time difference corresponding to two groups of associated moments in a pairwise manner, use any group of associated moments as the prefeature and the other group of associated moments as the postfeature, and confirm the feature difference according to the time difference between the prefeature and the postfeature; Select or lock the plane in which the feature difference is equal to the time difference from multiple vertical planes, that is, the first selected plane.
4. The method for optimizing voice interaction based on cross positioning according to claim 3, wherein Further including: Use the intersection line formed by any two of the first selected plane, the second selected plane, and the third selected plane as the spatial feature of the relevant voice signal. Specifically, use the spatial position information where the formed intersection line is located as the spatial feature of the relevant voice signal, where The spatial position information includes the length of the intersection line, the angle formed by the intersection line and the vertical direction, the angle formed by the intersection line and the horizontal direction perpendicular to the vertical direction, and the distance between the intersection line and the specified point.
5. The voice interaction optimization method based on cross positioning according to claim 1, characterized in that, The step of confirming the feature output text includes: Receive multiple groups of voice signals belonging to the same spatial feature, and based on the audio signal waveforms respectively associated with different voice signals, confirm the audio features associated with the corresponding voice signals; The following expression is used to calculate the audio features of the voice signal: TT = JJ × C1 + YY × C2; where TT represents the audio features of the voice signal; JJ represents the average value of several groups of horizontal vertical distances confirmed in the current audio signal waveform; YY represents the audio average value of the current audio signal confirmed by averaging several groups of confirmed audio frequencies; C1 and C2 are the first fixed coefficient factor and the second fixed coefficient factor respectively, the value of C1 ranges from 0.5 to 0.7, and the value of C2 ranges from 0.4 to 0.
5.
6. The voice interaction optimization method based on cross positioning according to claim 1, characterized in that It further includes: The following expression is used to calculate the comprehensive proportion of each text character and confirm the first group proportion: ZB1 = G1 ÷ G2; where ZB1 represents the first group proportion value calculated and confirmed, G1 represents the number of text characters with the same header and index text in the recognition feature output text; G2 represents the total number of text characters within the corresponding header recorded for the text characters with the same header and index text in the recognition feature output text; The following expression is used to calculate the comprehensive proportion of each text character and confirm the second group proportion: ZB2 = G1 ÷ G3; where ZB2 represents the second group proportion value calculated and confirmed; G1 represents the number of text characters with the same header and index text in the recognition feature output text; G3 represents the total number of index texts recorded for the text characters with the same header and index text in the recognition feature output text; The first group proportion value and the second group proportion value are averaged to determine the comprehensive proportion associated with the same text characters; If there are headers with multiple groups of the same text characters, the header associated with the maximum comprehensive proportion is selected from the comprehensive proportions associated with each different header and recorded as the selected header, and the selected header is used as the feature header.
7. The voice interaction optimization method based on cross positioning according to claim 1, characterized in that It further includes: The multiple groups of audio features are averaged to lock the audio feature average value, and the fluctuation range [YJ - Y2, YJ + Y2] is confirmed based on a preset value, where YJ is the locked audio feature average value; Y2 is the preset value, and the value of Y2 ranges from 2 to 4; Based on the fluctuation range, the fluctuation range with the largest total number of audio features is recorded as the determined range, and the voice signal associated with the abnormal feature is excluded using the determined range.
8. An optimized device for voice interaction based on cross positioning, characterized in that, It executes the voice interaction optimization method based on cross positioning according to any one of claims 1 to 7, and the voice interaction optimization device includes: An acquisition and processing module for acquiring voice signals of at least three acquisition nodes, confirming the spatial position information associated with different voice signals to obtain the spatial features associated with each voice signal, specifically including: confirming the time difference corresponding to two groups of associated moments in pairs, determining the feature plane, and constructing multiple vertical planes perpendicular to the feature plane; based on the feature plane and multiple vertical planes, confirming the feature difference to determine the selected plane; using the intersection line formed by any two selected planes as the spatial feature of the relevant voice signal; The confirmation module processes multiple groups of voice signals belonging to the same spatial feature based on the spatial features associated with the obtained voice signals, and then outputs text for the voice signals belonging to the same spatial feature to confirm the feature output text. The extraction and processing module extracts the stored content associated with each feature output text based on the self-built cloud database according to the feature output text generated by each spatial feature, and compiles each feature output text into an index text. The query and determination module, when receiving user input, returns the associated requested data content based on the index text recognized from the user input.
9. The voice interaction optimization device based on cross positioning according to claim 8, wherein It includes: Based on the feature plane and multiple vertical planes, select points to be used as the voice points to be confirmed, identify the distances between the voice points to be generated and the first feature point, and the distances between the voice points to be generated and the second feature point, so as to calculate the first time feature value of the acquisition node corresponding to the first feature point and the second time feature value of the acquisition node corresponding to the second feature point, and confirm the feature difference value to obtain the first selected plane determined by the first acquisition node and the second acquisition node. Repeat the relevant calculation and determination steps to obtain the second selected plane determined by the first acquisition node and the third acquisition node, and the third selected plane determined by the second acquisition node and the third acquisition node.
10. The voice interaction optimization device based on cross positioning according to claim 9, characterized in that, It includes: Calculate the time feature difference value according to the first time feature difference value and the second time feature value. Confirm the time difference values corresponding to two groups of associated moments in pairs, use any group of associated moments as the pre-feature and the other group of associated moments as the post-feature, and confirm the feature difference value according to the time difference value between the pre-feature and the post-feature. Select or lock the plane where the feature difference value is equal to the time difference value from multiple vertical planes, that is, the first selected plane.
Citation Information
Patent Citations
Equipment control method and device, storage medium and electronic device
CN113450798A
Display method and device, voice equipment and storage medium
CN115705849A
Human-computer interaction system and method based on auditory perception
CN118098228A
Adaptive inter-channel time difference estimation
CN119895493A
Robot and voice recognition equipment
CN218473315U