Reception process recording method and related device, server, system, and storage medium

By deploying mesh microphones in the reception area and using audio signal feature similarity and voiceprint matching for sound source localization, the problem of consultants needing to wear sensors was solved, enabling convenient recording and efficient analysis of the reception process.

CN116320886BActive Publication Date: 2026-07-24IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-12-07
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, consultants need to wear multiple sensors to collect data during customer interactions. This makes the accuracy and completeness of the data collection dependent on proper wearing and operation, and the lack of customer location information affects the value of the analysis.

Method used

Microphones are arranged in a mesh pattern in the reception area. The microphones pick up sound in a directional manner, and the target area and the person speaking are determined based on the similarity of audio signal features and voiceprint matching. The location of the customer and consultant is obtained by combining sound source localization, and the reception process is recorded.

Benefits of technology

The reception process can be recorded without the need for consultants to wear sensors, improving the convenience of recording and the value of analysis, and can simultaneously locate the activity positions of both clients and consultants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320886B_ABST
    Figure CN116320886B_ABST
Patent Text Reader

Abstract

The application discloses a reception process recording method and related device, server, system and storage medium, wherein the reception process recording method comprises: acquiring audio signals collected by each microphone in a mesh distribution in a reception place; in response to the feature similarity between audio signals oriented to the same grid region satisfying a first condition, determining that the grid region is a target region where a reception activity exists, and selecting the audio signal oriented to the target region as a target signal; performing voiceprint matching based on the target signal to determine a speaking object to which the target signal belongs; performing sound source positioning based on the collection time of each target signal belonging to the same speaking object to obtain an activity position of the speaking object in the reception place; and obtaining a reception record of the speaking object based on the activity position and the audio signal of the same speaking object at each moment in the reception process. The above scheme can improve the convenience of recording the reception process and improve the analysis value of the reception record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method for recording the reception process and related devices, servers, systems, and storage media. Background Technology

[0002] As people's living standards improve, when purchasing tangible and intangible goods, consumers not only consider the competitiveness of the product itself but also increasingly value the service experience during the purchase process. The service attitude, personal qualities, and communication skills of consultants during customer interactions all influence customers' purchasing experience and intentions. Therefore, service quality evaluation techniques during consultant interactions have gradually become a focus of attention.

[0003] In existing technologies, consultants typically wear various sensors, such as microphone switches, miniature motion sensors, and miniature positioning devices, to record the reception process. However, the accuracy and completeness of data collection in this method depend on the consultant correctly wearing and operating the sensors, which is quite inconvenient. Furthermore, since only the consultant's location information is recorded, not the customer's, the analytical value is significantly reduced. Therefore, improving the convenience of recording the reception process and enhancing the analytical value of the reception records have become urgent problems to be solved. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a method for recording the reception process, as well as related devices, servers, systems, and storage media, which can improve the convenience of recording the reception process and enhance the analytical value of the reception records.

[0005] To address the aforementioned technical problems, the first aspect of this application provides a method for recording a reception process, comprising: acquiring audio signals collected by various microphones distributed in a mesh pattern in a reception area; wherein the microphones respectively pick up sound in a direction to each mesh area intersecting at the microphone's location; in response to the feature similarity between audio signals picked up in the same mesh area satisfying a first condition, determining the mesh area as a target area where reception activities exist, and selecting audio signals picked up in the same mesh area as target signals; performing voiceprint matching based on the target signals to determine the speaking object to which the target signals belong; wherein the speaking object includes at least one of a customer or an advisor; performing sound source localization based on the acquisition time of each target signal belonging to the same speaking object to obtain the speaking object's activity location in the reception area; and obtaining a reception record of the speaking object based on the activity location and audio signals of the same speaking object at various times during the reception process.

[0006] To address the aforementioned technical problems, the first aspect of this application provides a reception process recording device, comprising: an acquisition module, a selection module, a determination module, a positioning module, and a recording module. The acquisition module is used to acquire audio signals collected by various microphones distributed in a mesh pattern within the reception area; wherein the microphones respectively pick up sound in a direction to each mesh area intersecting at the microphone's location. The selection module is used to determine a mesh area as a target area where reception activities exist, and select audio signals picked up in a direction to the target area as target signals, in response to a first condition being met by the feature similarity between audio signals picked up in the same mesh area. The determination module is used to perform voiceprint matching based on the target signals to determine the speaking object to which the target signals belong; wherein the speaking object includes at least one of a customer or an advisor. The positioning module is used to perform sound source localization based on the acquisition time of each target signal belonging to the same speaking object, obtaining the speaking object's activity location in the reception area. The recording module is used to obtain a reception record of the speaking object based on the activity location and audio signals of the same speaking object at various times during the reception process.

[0007] To address the aforementioned technical problems, a third aspect of this application provides a server comprising a communication circuit, a memory, and a processor. The communication circuit and the memory are respectively coupled to the processor. The communication circuit is used to acquire audio signals collected by a microphone. The memory stores program instructions, and the processor is used to execute the program instructions to implement the reception process recording method described in the first aspect.

[0008] To address the aforementioned technical problems, the fourth aspect of this application provides a reception process recording system, including a plurality of microphones distributed in a mesh pattern in the reception area and the server mentioned in the third aspect above, wherein the server is communicatively connected to the plurality of microphones to acquire audio signals collected by the microphones.

[0009] To address the aforementioned technical problems, the fifth aspect of this application provides a computer-readable storage medium storing program instructions executable by a processor, the program instructions being used to implement the reception process recording method of the first aspect described above.

[0010] The above solution acquires audio signals collected by each microphone in a mesh-like distribution within the reception area. Each microphone picks up sound directionally to a grid area intersecting with its location. Based on the condition that the feature similarity between audio signals picked up directionally to the same grid area satisfies a first condition, the grid area is determined as the target area where reception activities occur. Audio signals picked up directionally to the target area are selected as target signals. Voiceprint matching is performed based on the target signals to determine the speaker to whom the target signal belongs. The speaker includes at least one of the client or consultant. Sound source localization is then performed based on the acquisition time of each target signal belonging to the same speaker to obtain the speaker's location within the reception area. Finally, based on the speaker's location and audio signals at various moments during the reception process, a reception record of the speaker is obtained. Therefore, only a mesh-like distribution of microphones is needed to record the reception process, eliminating the need for consultants to wear multiple sensors such as microphone switches, miniature motion sensors, and miniature positioning devices. Furthermore, since sound source localization is performed based on the acquisition time of various target signals belonging to the same speaker during the reception process, the activity location of the speaker in the reception area can be obtained. Therefore, as long as the customer or consultant speaks during the reception, their activity location can be located simultaneously. This improves the convenience of recording the reception process and enhances the analytical value of the reception records. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating an embodiment of the reception process recording method of this application;

[0012] Figure 2 This is a schematic diagram of one embodiment of the microphone distribution method;

[0013] Figure 3 This is a schematic diagram of one embodiment of the target area;

[0014] Figure 4 This is a schematic diagram of an embodiment of the acquisition time of each microphone for directional sound pickup in a target area;

[0015] Figure 5 This is a schematic diagram of an embodiment of the activity positions of each speaking object at different times;

[0016] Figure 6 This is a schematic diagram of the framework of an embodiment of the reception process recording device of this application;

[0017] Figure 7 This is a schematic diagram of the framework of an embodiment of the server in this application;

[0018] Figure 8 This is a schematic diagram of the framework of an embodiment of the reception process recording system of this application;

[0019] Figure 9 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0020] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0021] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0022] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.

[0023] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the reception process recording method of this application. Specifically, it may include the following steps:

[0024] Step S11: Acquire the audio signals collected by each microphone in the mesh distribution in the reception area.

[0025] In one implementation scenario, the reception area can be set up according to the actual application needs. For example, in a car purchase scenario, the reception area could be a 4S store; or in a home purchase scenario, the reception area could be a marketing center. Other scenarios can be deduced similarly, and will not be listed here.

[0026] In one implementation scenario, the microphones in the reception area can be arranged in a checkerboard pattern. Please refer to the following: Figure 2 , Figure 2 This is a schematic diagram of one embodiment of the microphone distribution method. For example... Figure 2 As shown, solid circles filled with black represent microphones, arranged in a checkerboard pattern of four rows and five columns, forming a mesh-like distribution. Of course, Figure 2 The diagram shown is merely one possible distribution method in practical applications and does not limit the use of other microphone distribution methods. For example, microphones can also be distributed in a honeycomb mesh pattern, etc., which is not limited here.

[0027] In one implementation scenario, each microphone can be mounted on the ceiling of the reception area. In this case, the clarity of sound pickup can be improved by positioning the microphones horizontally downwards to capture as much sound as possible from below the ceiling.

[0028] In this embodiment of the disclosure, the microphones respectively pick up sound in a directional manner towards each grid area intersecting at the microphone's location. Using microphone b2 ( Figure 2 Taking the microphone in the second row and second column as an example, the grid areas intersecting its position include: grid area 1 at its upper left, grid area 2 at its upper right, grid area 3 at its lower right, and grid area 4 at its lower left. In other words, microphone b2 has four pickup areas, namely grid areas 1 to 4. Other cases can be deduced similarly, and will not be listed here.

[0029] Step S12: In response to the fact that the feature similarity between audio signals that are picked up in the same grid area satisfies the first condition, the grid area is determined to be the target area where reception activities exist, and the audio signal that is picked up in the target area is selected as the target signal.

[0030] In one implementation scenario, for each grid region, audio signals picked up directionally towards that grid region can be acquired, and the audio features of these audio signals can be extracted respectively, thereby allowing the calculation of the pairwise feature similarity between these audio features. Figure 2 Taking grid area 3 as an example, microphone b2 can pick up the audio signal of grid area 3 by directional pickup to the lower right, microphone c2 can pick up the audio signal of grid area 3 by directional pickup to the lower left, microphone c3 can pick up the audio signal of grid area 3 by directional pickup to the upper left, and microphone b3 can pick up the audio signal of grid area 3 by directional pickup to the upper right. Other cases can be deduced similarly, and will not be listed here. It should be noted that the terms "lower right," "lower left," "upper left," and "upper right" mentioned above are all based on... Figure 2 From the planar perspective shown. In practical applications, since the microphones mentioned above are all in three-dimensional space, therefore, if we consider... Figure 2 The upward direction represents the rear of the three-dimensional space, and the downward direction represents the front of the three-dimensional space. Therefore, "lower right" corresponds to "front right," "lower left" corresponds to "front left," "upper left" corresponds to "rear left," and "upper right" corresponds to "rear right." Other cases can be deduced similarly, and will not be listed here. In addition, the above audio features may include, but are not limited to, FBank, MFCC, etc., which are not limited here.

[0031] In one implementation scenario, the first condition may include feature similarity being higher than a similarity threshold. Specifically, the similarity threshold can be set according to the actual application requirements. For example, when the detection requirements are high, the similarity threshold can be set larger, while when the detection requirements are relatively lenient, the similarity threshold can be set appropriately smaller.

[0032] In one implementation scenario, if the feature similarity between any pair of audio signals picked up from the same grid area satisfies the first condition, the grid area can be determined as a target area where a reception activity is taking place. Please refer to [reference needed]. Figure 3 , Figure 3 This is a schematic diagram of one embodiment of the target area. For example... Figure 3 As shown, calculations show that the pairwise similarity of the audio signals picked up directionally towards grid area 3 satisfies the first condition. Therefore, grid area 3 can be determined as the target area where reception activities take place. For example, in grid area 3, white-filled circles represent customers, and gray-filled circles represent consultants. Based on this, audio signals picked up directionally towards grid area 3 can be selected as target signals. That is, the audio signals picked up directionally by microphone b2, microphone c2, microphone c3, and microphone b3 towards grid area 3 can each be used as target signals.

[0033] It should be noted that during a conversation, there is usually a gap, long or short, between the audio signals of different speakers. This gap can be used to distinguish each audio segment, and in subsequent processing, audio segments can be used as the basic unit to avoid interference caused by mixing audio from different speakers.

[0034] Step S13: Perform voiceprint matching based on the target signal to determine the speaker to whom the target signal belongs.

[0035] In this embodiment of the disclosure, the speaking object includes at least one of a customer or an advisor. As one possible implementation, the voiceprint features of different speaking objects can be pre-stored as pre-stored voiceprint features. For example, the voiceprint features of each advisor can be pre-stored, as can the voiceprint features of customers who have visited. Furthermore, each pre-stored voiceprint feature can be labeled with its corresponding speaking object. Based on this, the voiceprint features of the target signal can be extracted as the target voiceprint features, and the matching degree between the target voiceprint features and each pre-stored voiceprint feature can be obtained. For example, the matching degree between the target voiceprint features and the pre-stored voiceprint features can be calculated using methods such as cosine similarity or inner product. Based on this, in response to the matching degree between the pre-stored voiceprint features and the target voiceprint features satisfying a second condition, the speaking object to which the pre-stored voiceprint features belong can be used as the speaking object to which the target signal belongs. It should be noted that the second condition can also be set to a matching degree higher than a matching degree threshold. Similar to the first condition, the matching degree threshold is set according to the actual application requirements. For example, when high matching accuracy is required, the matching threshold can be set slightly higher; conversely, when matching accuracy requirements are relatively relaxed, the matching threshold can be set slightly lower. Furthermore, when multiple pre-stored voiceprint features have a matching degree higher than the target voiceprint feature, the speaker to which the pre-stored voiceprint feature with the highest matching degree belongs can be selected as the speaker to which the target signal belongs. Conversely, if no pre-stored voiceprint feature satisfies the second condition, a new speaker can be created for the target signal. More precisely, a new object ID can be created, and a visit record can be created, binding detailed object information (e.g., name, age, etc.) to the new object ID. The target voiceprint feature can also be further recorded as a pre-stored voiceprint feature of the new speaker for subsequent matching. The above method extracts the target voiceprint features of the target signal, obtains the matching degree between the target voiceprint features and each pre-stored voiceprint feature, and in response to the matching degree of the pre-stored voiceprint features satisfying the second condition, the speaking object to which the pre-stored voiceprint features belong is taken as the speaking object to which the target signal belongs. In response to the absence of a pre-stored voiceprint feature whose matching degree satisfies the second condition, a new speaking object is created for the target signal, and the target voiceprint features are recorded as the pre-stored voiceprint features of the new speaking object. Therefore, it is possible to determine the speaking object to which each segment of the target signal belongs based on the voiceprint, which helps to improve the accuracy of reception records.

[0036] Step S14: Based on the acquisition time of each target signal belonging to the same speaking object, perform sound source localization to obtain the activity location of the speaking object in the reception area.

[0037] In one implementation scenario, deep learning technology can be used to obtain an end-to-end sound source localization model (SSL). The target signal mentioned above is then input into the end-to-end sound source localization model for processing to obtain the location of the speaker in the reception area. Specifically, the end-to-end sound source localization model can employ techniques such as deep neural networks or convolutional neural networks, which are not limited here.

[0038] In one implementation scenario, unlike the aforementioned implementations, end-to-end sound source localization models may struggle to perform sound source localization tasks when labeled data is insufficient. In this case, a microphone corresponding to the earliest acquisition time of the target signal among speakers can be selected as the target microphone. A microphone directionally picking up sound towards the target area outside the target microphone can be used as a reference microphone. This allows us to obtain the time deviation between the target signal acquired by the reference microphone and the target signal acquired by the target microphone. Then, considering the active location and the propagation time of the target signal to the target microphone as unknowns, a system of equations is established based on the installation location, active location, time deviation, and propagation time of the microphone directionally picking up sound towards the target area. Solving this system of equations yields the active location of the speaker in the reception area. This method, by combining time deviation and spatial location to solve a system of equations to obtain the active location, helps simplify the complexity of sound source localization.

[0039] In a specific implementation scenario, please refer to the relevant documents. Figure 3 and Figure 4 , Figure 4 This is a schematic diagram of an embodiment of the acquisition time of each microphone for directional sound pickup in a target area. Taking grid area 3 as the target area as an example, the acquisition times of microphones b2, c2, b3, and c3 that acquire the target signal belonging to the same speaking object can be compared. For example, for the customer represented by the white-filled circle, microphone b2 acquires the target signal earliest, so microphone b2 can be used as the target microphone, and microphones c2, c3, and b3 as reference microphones. Other microphones can be deduced similarly, and will not be listed here. Further, the time deviations between the acquisition times of each reference microphone (i.e., microphones c2, c3, and b3) and the acquisition time of the target microphone (i.e., microphone b2) can be obtained, denoted as t1, t2, and t3 respectively. Based on this, the activity position and the propagation time of the target signal to the target microphone can be treated as unknowns, and a system of equations can be solved. For example, the activity position can be denoted as (x... cus y cus , z cus The propagation time is denoted as t. Furthermore, for the consultant represented by the gray-filled circle, the same principle applies; the solution is obtained by solving the simultaneous variance set, which will not be elaborated further here.

[0040] In a specific implementation scenario, when solving the system of equations for a target microphone, a first equation representing the distance between the speaker and the target microphone can be obtained based on the speaker's location and the microphone's installation location. A second equation representing the distance between the speaker and the target microphone can be obtained based on the speed and duration of sound propagation. Finally, the first equation can be established based on both equations. In other words, the first equation based on spatial location and the second equation based on sound propagation express the same meaning and can be connected by an equal sign to establish the first equation. Please refer to further details. Figure 3 and Figure 4 For the target microphone b2, the following first equation can be established:

[0041] (x b2 -x cus ) 2 +(y b2 -y cus ) 2 +(z b2 -z cus ) 2 =(t*v sound ) 2 ……(1)

[0042] In the above formula (1), (x b2 y b2 , z b2 ) indicates the installation location of the target microphone b2, V sound This represents the speed of sound. For example, the speed of sound in air can be denoted as 340 m / s. Furthermore, the left side of the equal sign represents the first equation, and the right side represents the second equation. Other cases can be deduced similarly to establish the first equation, which will not be illustrated further here.

[0043] In a specific implementation scenario, for a reference microphone, when setting up a system of equations, a third equation representing the distance between the speaker and the reference microphone can be obtained based on the speaker's location and the microphone's installation location. A fourth equation representing the distance between the speaker and the reference microphone can be obtained based on the speed of sound propagation, propagation time, and the corresponding time deviation of the reference microphone. Finally, a second equation can be established based on the third and fourth equations. In other words, the third equation based on spatial location and the fourth equation based on sound propagation express the same meaning and can be connected by an equal sign to establish the second equation. Please refer to further details. Figure 3 and Figure 4 For reference microphones c2, c3, and b3, the following second equation can be established:

[0044] (xc2 -x cus ) 2 +(y c2 -y cus ) 2 +(z c2 -z cus ) 2 =((t+t1)*v sound ) 2 …(2)

[0045] (x c3 -x cus ) 2 +(y c3 -y cus ) 2 +(z c3 -z cus ) 2 =((t+t2)*v sound ) 2 …(3)

[0046] (x b3 -x cus ) 2 +(y b3 -y cus ) 2 +(z b3 -z cus ) 2 =((t+t3)*v sound ) 2 …(4)

[0047] In formulas (2) to (3) above, (x c2 y c2 , z c2 ), (x c3 y c3 , z c3 ), (x b3 y b3 , z b3 The numbers () represent the mounting positions of reference microphones c2, c3, and b3, respectively. sound This represents the speed of sound propagation. Furthermore, the left side of the equals sign represents the third equation, and the right side represents the fourth equation. Other cases can be deduced similarly to establish the first equation; examples will not be provided here.

[0048] In a specific implementation scenario, further, after obtaining the first and second equations, a system of equations can be established based on the first and second equations. Please refer to the following references. Figure 3 and Figure 4Combining the equations shown in formulas (1) to (4) above, we obtain a system of equations, and by solving it, we can obtain the result. Figure 3 The white-filled circle indicates the customer's activity location in the reception area (x). cus y cus , z cus Other cases can be deduced similarly, and will not be listed here. In the above method, for the target microphone, based on the activity position and the installation position of the target microphone, a first equation representing the distance between the speaker and the target microphone is obtained. Based on the speed of sound propagation and propagation time, a second equation representing the distance between the speaker and the target microphone is obtained. Based on the first and second equations, a first equation is established. For the reference microphone, based on the activity position and the installation position of the reference microphone, a third equation representing the distance between the speaker and the reference microphone is obtained. Based on the speed of sound propagation, propagation time, and the corresponding time deviation of the reference microphone, a fourth equation representing the distance between the speaker and the reference microphone is obtained. Based on the third and fourth equations, a second equation is established. Therefore, by simultaneously solving the first and second equations, a system of equations can be established, thus distinguishing between the target microphone and the reference microphone. The same principle is used to establish equations for both, and by simultaneously solving the system of equations, the activity position of the speaker can be obtained, which helps improve the accuracy of the activity position.

[0049] It should be noted that after obtaining the location of the person being spoken to, the earliest collection time of that person can be bound to that location, meaning that at that earliest collection time, the person being spoken to was at that location in the reception area.

[0050] Step S15: Based on the activity location and audio signal of the same person speaking at various times during the reception process, obtain the reception record of the person speaking.

[0051] Specifically, for each person being spoken to, their activity location and audio signals at various moments during the reception process can be sorted chronologically to obtain the reception record for that person.

[0052] In one implementation scenario, after obtaining the reception records, the probability of each consultant serving the customer is calculated based on the consultant whose activity location is closest to the customer's location at each time point. Then, based on this probability, the consultant serving the customer is determined. Finally, a reception recording is obtained based on the audio signals of the customer and their consultant during the reception process. For example, the audio signals of the customer and their consultant during the reception process can be sorted chronologically to obtain the customer's reception recording. This method, by calculating the probability of each customer being served by different consultants and then determining the consultant based on these probabilities, allows for the accurate creation of reception recordings for each customer.

[0053] In a specific implementation scenario, to calculate the above probabilities, we can determine the number of times each consultant is closest to the customer at each time point, based on the consultant whose activity location is closest to the customer's at each time point. Then, based on the number of times each consultant is closest to the customer, we can obtain the probability that the customer will be served by each consultant. For example, we can denote the total number of the aforementioned time points as N, and the number of times the i-th consultant is closest to the j-th customer at each time point as M(i, j). Then, the probability P(i, j) that the j-th customer will be served by the i-th consultant can be expressed as:

[0054] P(i,j)=M(i,j) / N......(5)

[0055] In a specific implementation scenario, please refer to the relevant documents. Figure 5 , Figure 5 This is a schematic diagram illustrating an embodiment of the activity positions of various speaking objects at different times. For example... Figure 5 As shown, from time t1 to t3, for customer 1, the consultants closest to customer 1 are consultant 1, consultant 3, and consultant 1, respectively. That is, from time t1 to t3, consultant 1 is closest to customer 1 2 times, consultant 2 is closest to customer 1 0 times, and consultant 3 is closest to customer 1 1 time. Therefore, the probability that customer 1 is served by consultant 1 is 2 / 3, the probability of being served by consultant 2 is 0, and the probability of being served by consultant 3 is 1 / 3. Based on this, the consultant with the highest probability can be selected as the consultant serving the customer. In other words, in this case, it can be determined that the consultant serving customer 1 is consultant 1. Other cases can be deduced similarly, and will not be listed here. The above method, based on the consultant whose activity position is closest to the customer's activity position at each time point, obtains the number of times each consultant is closest to the customer, and then, based on the number of times each consultant is closest to the customer, obtains the probability that the customer is served by each consultant, which improves the accuracy of probability calculation.

[0056] In one implementation scenario, the reception area can be pre-divided into several functional areas. For example, taking a 4S dealership as an example, it can be divided into functional areas including, but not limited to: a negotiation area, a vehicle delivery area, a display area for vehicle A, and a display area for vehicle B. Other cases can be deduced similarly, and will not be listed here. After obtaining the reception records, the activity positions of the same person during the reception can be traversed chronologically, recording the times when the activity position first enters and first leaves the same functional area. Based on these times, the duration of the person's stay in the corresponding functional area can be obtained. Specifically, the duration of the person's stay in the corresponding functional area can be obtained by subtracting the times when they first enter and first leave the same functional area. For example, for the i-th customer, the times when they first enter and first leave the same functional area j are respectively... and The duration of customer i's stay in functional area j can be determined by... minus The calculation is as follows. Other cases can be deduced similarly, and will not be listed here. The above method, in chronological order, traverses the activity positions of the same speaker at various moments during the reception process, and records the moment when the activity position first falls into and first leaves the same functional area. Based on the moment when the activity position first falls into and first leaves the same functional area, the duration of the speaker's stay in the corresponding functional area can be obtained, which can improve the accuracy of the speaker's stay analysis.

[0057] In a specific implementation scenario, as mentioned earlier, the consultant responsible for a customer can be determined based on the probability that the customer is received by each consultant. Specifically, during the same reception process, the activity locations of the customer and their assigned consultant can be summarized, sorted, and traversed according to time sequence, and iterated to check whether the activity locations fall within the functional area. It should be noted that if the consultant does not have a corresponding audio signal at a specific interval (e.g., 5 minutes, 10 minutes, etc.), the current reception can be considered to have ended.

[0058] In a specific implementation scenario, as mentioned above, the customer reception recording can also be obtained based on the audio signals of the customer and their reception consultant during the reception process. In order to establish a detailed reception file for each customer, the customer reception recording and the duration of stay in each functional area can be saved as the customer's reception file.

[0059] The above solution acquires audio signals collected by each microphone in a mesh-like distribution within the reception area. Each microphone picks up sound directionally to a grid area intersecting with its location. Based on the condition that the feature similarity between audio signals picked up directionally to the same grid area satisfies a first condition, the grid area is determined as the target area where reception activities occur. Audio signals picked up directionally to the target area are selected as target signals. Voiceprint matching is performed based on the target signals to determine the speaker to whom the target signal belongs. The speaker includes at least one of the client or consultant. Sound source localization is then performed based on the acquisition time of each target signal belonging to the same speaker to obtain the speaker's location within the reception area. Finally, based on the speaker's location and audio signals at various moments during the reception process, a reception record of the speaker is obtained. Therefore, only a mesh-like distribution of microphones is needed to record the reception process, eliminating the need for consultants to wear multiple sensors such as microphone switches, miniature motion sensors, and miniature positioning devices. Furthermore, since sound source localization is performed based on the acquisition time of various target signals belonging to the same speaker during the reception process, the activity location of the speaker in the reception area can be obtained. Therefore, as long as the customer or consultant speaks during the reception, their activity location can be located simultaneously. This improves the convenience of recording the reception process and enhances the analytical value of the reception records.

[0060] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the reception process recording device 60 of this application. The reception process recording device 60 includes: an acquisition module 61, a selection module 62, a determination module 63, a positioning module 64, and a recording module 65. The acquisition module 61 is used to acquire audio signals collected by each microphone distributed in a mesh pattern in the reception area; wherein, the microphones respectively pick up sound in a direction to each mesh area intersecting at the microphone location; the selection module 62 is used to determine the mesh area as a target area where reception activities exist in response to the feature similarity between audio signals picked up in the same mesh area satisfying a first condition, and select the audio signals picked up in the target area as target signals; the determination module 63 is used to perform voiceprint matching based on the target signals to determine the speaking object to which the target signals belong; wherein, the speaking object includes at least one of a customer and an advisor; the positioning module 64 is used to perform sound source localization based on the acquisition time of each target signal belonging to the same speaking object to obtain the activity position of the speaking object in the reception area; the recording module 65 is used to obtain the reception record of the speaking object based on the activity position and audio signals of the same speaking object at various times during the reception process.

[0061] The above solution, because the reception process recording device 60 can implement the steps in the above-described reception process recording method embodiment, only requires the arrangement of a mesh of microphones in the reception area to record the reception process, without requiring consultants to wear multiple sensors such as microphone switches, miniature motion sensors, and miniature positioning devices. Furthermore, since sound source localization is performed based on the acquisition time of signals from various targets belonging to the same speaker during the reception process, the activity position of the speaker in the reception area can be obtained. Therefore, as long as the customer or consultant speaks during the reception process, their activity position can be located simultaneously. Thus, the convenience of recording the reception process is improved, and the analytical value of the reception records is enhanced.

[0062] In some disclosed embodiments, the positioning module 64 includes a microphone selection submodule, used to select the microphone corresponding to the target signal with the earliest acquisition time among the same speaking objects as the target microphone, and to use the microphone that picks up sound directionally towards the target area outside the target microphone as the reference microphone; the positioning module 64 includes a time deviation acquisition submodule, used to acquire the time deviation between the target signal acquired by the reference microphone and the target signal acquired by the target microphone; the positioning module 64 includes an equation system simultaneous solution submodule, used to use the activity position and the propagation time of the target signal to the target microphone as unknowns, and to simultaneously solve a system of equations based on the installation position, activity position, time deviation and propagation time of the microphone that picks up sound directionally towards the target area; the positioning module 64 includes an activity position solving submodule, used to solve the system of equations to obtain the activity position of the speaking object in the reception area.

[0063] Therefore, solving the simultaneous equations by combining time deviation and spatial location to obtain the location of the activity helps to simplify the complexity of sound source localization.

[0064] In some disclosed embodiments, the equation system simultaneous submodule includes a first establishment unit, used to, for a target microphone, derive a first equation representing the distance between the speaking object and the target microphone based on the active position and the installation position of the target microphone, and derive a second equation representing the distance between the speaking object and the target microphone based on the sound propagation speed and propagation time, and establish a first equation based on the first and second equations; the equation system simultaneous submodule includes a second establishment unit, used to, for a reference microphone, derive a third equation representing the distance between the speaking object and the reference microphone based on the active position and the installation position of the reference microphone, and derive a fourth equation representing the distance between the speaking object and the reference microphone based on the sound propagation speed, propagation time and the time deviation corresponding to the reference microphone, and establish a second equation based on the third and fourth equations; the equation system simultaneous submodule includes an equation simultaneous unit, used to simultaneously establish a system of equations based on the first and second equations.

[0065] Therefore, being able to distinguish between the target microphone and the reference microphone by establishing equations using the same principle, and then solving the system of equations to obtain the location of the speaker, helps to improve the accuracy of the location.

[0066] In some disclosed embodiments, the reception process recording device 60 further includes a probability statistics module, used to calculate the probability that the customer will be received by each consultant based on the consultant whose activity location is closest to the customer's activity location at each time point; the reception process recording device 60 also includes a consultant determination module, used to determine the customer's receiving consultant based on the probability that the customer will be received by each consultant; the reception process recording device 60 also includes a recording acquisition module, used to obtain the customer's reception recording based on the audio signals of the customer and their receiving consultant during the reception process.

[0067] Therefore, by statistically analyzing the probability of each customer being received by a different consultant, and then determining the consultant who will receive the customer based on the above probability, the audio signals of the customer and their consultant during the reception process can be used to obtain the customer's reception recordings, thereby enabling the accurate creation of reception recordings for each customer.

[0068] In some disclosed embodiments, the probability statistics module includes a frequency acquisition submodule, used to obtain the number of times each consultant is closest to the customer based on the consultant whose activity location is closest to the customer's activity location at each time point; the probability statistics module includes a probability calculation submodule, used to obtain the probability that the customer will be received by each consultant based on the number of times each consultant is closest to the customer.

[0069] Therefore, by identifying the consultant whose activity location is closest to the customer's activity location at each time point, we can obtain the number of times each consultant is closest to the customer. Then, based on the number of times each consultant is closest to the customer, we can obtain the probability that the customer will be served by each consultant, which can improve the accuracy of probability calculation.

[0070] In some disclosed embodiments, the reception area is pre-divided into several functional areas. The process recording device 60 also includes a position traversal module, which is used to traverse the activity positions of the same speaker at various moments during the reception process in chronological order. The process recording device 60 also includes a time recording module, which is used to record the moment when the activity position first falls into the same functional area and the moment when it first falls out of the same functional area. The process recording device 60 also includes a duration calculation module, which is used to obtain the duration of the speaker's stay in the corresponding functional area based on the moment when the position first falls into the same functional area and the moment when it first falls out of the same functional area.

[0071] Therefore, by traversing the activity positions of the same speaker at various moments during the reception process in chronological order, and recording the first time the activity position falls into and the first time it leaves the same functional area, the duration of the speaker's stay in the corresponding functional area can be obtained based on the first time the activity position falls into and the first time it leaves the same functional area. This can improve the accuracy of the speaker's stay analysis.

[0072] In some disclosed embodiments, the determining module 63 includes a voiceprint extraction submodule for extracting target voiceprint features of the target signal; the determining module 63 includes a matching degree calculation submodule for obtaining the matching degree between the target voiceprint features and each pre-stored voiceprint feature; the determining module 63 includes a first response submodule for, in response to the matching degree of the pre-stored voiceprint features satisfying a second condition, taking the speaking object to which the pre-stored voiceprint features belong as the speaking object to which the target signal belongs; the determining module 63 includes a second response submodule for, in response to the absence of a pre-stored voiceprint feature whose matching degree satisfies the second condition, creating a new speaking object for the target signal and recording the target voiceprint features as the pre-stored voiceprint features of the new speaking object.

[0073] Therefore, by extracting the target voiceprint features of the target signal, obtaining the matching degree between the target voiceprint features and each pre-stored voiceprint feature, and in response to the matching degree of the pre-stored voiceprint features satisfying the second condition, the speaking object to which the pre-stored voiceprint features belong is taken as the speaking object to which the target signal belongs. In response to the absence of a pre-stored voiceprint feature whose matching degree satisfies the second condition, a new speaking object is created for the target signal, and the target voiceprint features are recorded as the pre-stored voiceprint features of the new speaking object. Thus, it is possible to determine the speaking object to which each segment of the target signal belongs based on the voiceprint, which helps to improve the accuracy of reception records.

[0074] In some disclosed embodiments, the microphones in the reception area are arranged in a checkerboard pattern; and / or, the microphones are mounted on the ceiling of the reception area.

[0075] Therefore, arranging the microphones in the reception area in a checkerboard pattern helps reduce the complexity of microphone placement; and installing the microphones on the ceiling of the reception area helps to minimize noise pickup and improve audio acquisition quality.

[0076] Please see Figure 7 , Figure 7 This is a schematic diagram of a framework of an embodiment of the server 70 of this application. The server 70 includes a communication circuit 71, a memory 72, and a processor 73. The memory 72 stores program instructions, and the processor 73 is used to execute the program instructions to implement the steps in any of the above-described embodiments of the reception process recording method.

[0077] Specifically, processor 73 controls itself, as well as communication circuit 71 and memory 72, to implement the steps in any of the above-described reception process recording method embodiments. Processor 73 can also be referred to as a CPU (Central Processing Unit). Processor 73 may be an integrated circuit chip with signal processing capabilities. Processor 73 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 73 can be implemented using integrated circuit chips.

[0078] The above solution, because the server 70 can implement the steps in the above-described reception process recording method embodiment, only requires the placement of a mesh of microphones in the reception area to record the reception process, without requiring consultants to wear multiple sensors such as microphone switches, miniature motion sensors, and miniature positioning devices. Furthermore, since sound source localization is performed based on the acquisition time of signals from various targets belonging to the same speaker during the reception process, the activity position of the speaker in the reception area can be obtained. Therefore, as long as the customer or consultant speaks during the reception process, their activity position can be located simultaneously. Thus, the convenience of recording the reception process is improved, and the analytical value of the reception records is enhanced.

[0079] Please see Figure 8 , Figure 8 This is a schematic diagram of a framework of an embodiment of the reception process recording system 80 of this application. The reception process recording system 80 includes a plurality of microphones 81 distributed in a mesh pattern in the reception area and a server 82 as described in the aforementioned disclosed embodiments. The server 82 is communicatively connected to the plurality of microphones 81 to acquire audio signals collected by the microphones 81. It should be noted that the microphones 81 can communicate with the server 82 wirelessly (such as Bluetooth, Zigbee, etc.) or wiredly (such as network cable, fiber optic, etc.), and no limitation is made here.

[0080] In the above-described solution, since the server 82 in the reception process recording system 80 is the same server shown in the server embodiment, only a mesh-distributed array of microphones needs to be arranged in the reception area to record the reception process. This eliminates the need for consultants to wear multiple sensors such as microphone switches, miniature motion sensors, and miniature positioning devices. Furthermore, because sound source localization is performed based on the acquisition time of signals from various targets belonging to the same speaker during the reception process, the activity location of the speaker in the reception area can be determined. Therefore, as long as the customer or consultant speaks during the reception process, their activity location can be located simultaneously. This improves the convenience of recording the reception process and enhances the analytical value of the reception records.

[0081] Please see Figure 9 , Figure 9 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 90 of this application. The computer-readable storage medium 90 stores program instructions 91 that can be executed by a processor. The program instructions 91 are used to implement the steps in any of the above-described embodiments of the reception process recording method.

[0082] The above solution, because the computer-readable storage medium 90 can implement the steps in the above-described reception process recording method embodiment, only requires the arrangement of a mesh of microphones in the reception area to record the reception process, without requiring consultants to wear multiple sensors such as microphone switches, miniature motion sensors, and miniature positioning devices. Furthermore, since sound source localization is performed based on the acquisition time of signals from various targets belonging to the same speaker during the reception process, the activity position of the speaker in the reception area can be obtained. Therefore, as long as the customer or consultant speaks during the reception process, their activity position can be located simultaneously. Thus, the convenience of recording the reception process is improved, and the analytical value of the reception records is enhanced.

[0083] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0084] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0085] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0086] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0087] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0088] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0089] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A method for recording the reception process, characterized in that, include: The system acquires audio signals collected by each microphone in a grid-like distribution within the reception area; wherein each microphone picks up sound directionally in each grid area intersecting at its location. In response to the fact that the feature similarity between audio signals that are directionally picked up in the same grid area satisfies a first condition, the grid area is determined to be a target area where reception activities exist, and the audio signal that is directionally picked up in the target area is selected as the target signal; wherein, the microphones in the reception area are distributed in a checkerboard pattern. Voiceprint matching is performed based on the target signal to determine the speaking party to which the target signal belongs; wherein, the speaking party includes at least one of a customer or an advisor; Select the microphone corresponding to the target signal with the earliest acquisition time among the same speaking objects as the target microphone, and use the microphones that pick up sound directionally from the target microphone toward the target area as reference microphones; The time deviation between the target signal acquired by the reference microphone and the target signal acquired by the target microphone is obtained; Using the activity location and the propagation time of the target signal to the target microphone as unknowns, a system of equations is established based on the installation location of the microphone that picks up sound directionally to the target area, the activity location, the time deviation, and the propagation time. Solving the system of equations yields the location of the speaker in the reception area. Based on the activity location and audio signal of the same person speaking at various moments during the reception process, a reception record of the person speaking is obtained.

2. The method according to claim 1, characterized in that, The system of equations, based on the installation location of the microphone directionally picking up sound towards the target area, the active location, the time deviation, and the propagation time, includes: For the target microphone, based on the activity position and the installation position of the target microphone, a first formula representing the distance between the speaking object and the target microphone is obtained, and based on the speed of sound propagation and the propagation time, a second formula representing the distance between the speaking object and the target microphone is obtained, and based on the first formula and the second formula, a first equation is established; For the reference microphone, based on the activity position and the installation position of the reference microphone, a third equation is obtained to characterize the distance between the speaking object and the reference microphone. Based on the speed of sound propagation, the propagation time, and the time deviation corresponding to the reference microphone, a fourth equation is obtained to characterize the distance between the speaking object and the reference microphone. Based on the third equation and the fourth equation, a second equation is established. Based on the first equation and the second equation, the system of equations is established simultaneously.

3. The method according to claim 1, characterized in that, After obtaining the reception record of the speaking object based on the activity location and audio signal of the same speaking object at various times during the reception process, the method further includes: Based on the consultant whose activity location is closest to the customer's activity location at each of the aforementioned times, the probability that the customer will be served by each of the aforementioned consultants is calculated. The consultant who will receive the customer is determined based on the probability that the customer will be received by each of the consultants. Based on the audio signals of the customer and their reception consultant during the reception process, a reception recording of the customer is obtained.

4. The method according to claim 3, characterized in that, The method of calculating the probability that a customer will be served by each consultant based on the consultant whose activity location is closest to the customer's activity location at each of the aforementioned times includes: Based on the consultant whose activity location is closest to the customer's activity location at each of the aforementioned times, the number of times each consultant is closest to the customer is obtained; Based on the number of times each consultant is closest to the customer, the probability that the customer will be served by each of the consultants is obtained.

5. The method according to claim 1, characterized in that, The reception area is pre-divided into several functional areas. After obtaining the reception record of the speaker based on the speaker's activity position and audio signal at various times during the reception process, the method further includes: According to the chronological order, traverse the activity positions of the same speaking object at various moments during the reception process, and record the moment when the activity position first falls into and the moment when it first leaves the same functional area; The duration of the speaking object's stay in the corresponding functional area is obtained based on the times when it first falls into and first leaves the same functional area.

6. The method according to claim 1, characterized in that, The step of performing voiceprint matching based on the target signal to determine the speaking party to which the target signal belongs includes: Extract the target voiceprint features of the target signal; Obtain the matching degree between the target voiceprint features and each pre-stored voiceprint feature; In response to the matching degree of the pre-stored voiceprint features satisfying the second condition, the speaking object to which the pre-stored voiceprint features belong is taken as the speaking object to which the target signal belongs. In response to the absence of a pre-stored voiceprint feature whose matching degree satisfies the second condition, a new speaking object is created for the target signal, and the target voiceprint feature is recorded as the pre-stored voiceprint feature of the new speaking object.

7. The method according to any one of claims 1 to 6, characterized in that, Each microphone is mounted on the ceiling of the reception area.

8. A reception process recording device, characterized in that, include: The acquisition module is used to acquire audio signals collected by each microphone in the mesh-distributed reception area; wherein, each microphone picks up sound directionally to each mesh area intersecting at the location of the microphone; The selection module is configured to determine the grid area as a target area where reception activities exist in response to a first condition being met by the feature similarity between audio signals that are directionally picked up towards the same grid area, and to select the audio signals that are directionally picked up towards the target area as target signals; wherein, the microphones in the reception area are distributed in a checkerboard pattern. The determination module is used to perform voiceprint matching based on the target signal to determine the speaking object to which the target signal belongs; wherein the speaking object includes at least one of a customer or an advisor; The microphone selection submodule is used to select the microphone corresponding to the target signal with the earliest acquisition time among the same speaking objects as the target microphone, and to use the microphones that pick up sound directionally to the target area outside the target microphone as reference microphones. The time deviation acquisition submodule is used to acquire the time deviation between the target signal acquired by the reference microphone and the target signal acquired by the target microphone; The equation system submodule is used to simultaneously solve equations based on the installation position of the microphone that picks up sound in the target area, the active position, the time deviation, and the propagation time, taking the active position and the propagation time of the target signal as unknowns. The activity location solution submodule is used to solve the system of equations to obtain the activity location of the speaking object in the reception area; The recording module is used to obtain a reception record of the speaking object based on the activity position and audio signal of the same speaking object at various times during the reception process.

9. A server, characterized in that, The device includes a communication circuit, a memory, and a processor. The communication circuit and the memory are respectively coupled to the processor. The communication circuit is used to acquire audio signals collected by a microphone. The memory stores program instructions. The processor is used to execute the program instructions to implement the reception process recording method according to any one of claims 1 to 7.

10. A reception process recording system, characterized in that, The system includes a plurality of microphones distributed in a mesh pattern in the reception area and a server as described in claim 9, wherein the server is communicatively connected to the plurality of microphones to acquire audio signals collected by the microphones.

11. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the reception process recording method according to any one of claims 1 to 7.