AI-based intelligent chest badge collection and analysis method and system
By dynamically adjusting the microphone array sensitivity using a large AI-based model and structured light module, the accuracy issues of existing AI quality inspection systems in semantic understanding and noisy environments are solved, achieving higher service quality judgment and operational efficiency.
Patent Information
- Application Number
- CN202511248411.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing AI quality inspection systems rely on keyword matching, which cannot understand contextual semantics and implicit information, leading to missed detections and false alarms. Furthermore, the audio quality is poor in noisy environments, affecting the accuracy of service quality judgment.
The system employs a large AI-based model for audio transcription and text analysis, combines structured light modules to identify personnel location and facial information, dynamically adjusts the microphone array sensitivity, and uses a noise recognition strategy to shut down noisy microphones, thereby improving the clarity of audio acquisition.
It improves the ability to understand semantics and hidden information, enhances the accuracy of service quality judgment, reduces the impact of noise, and improves the accuracy and operational efficiency of the quality inspection system.
Smart Images

Figure CN120748448B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of AI large model analysis, in particular to an AI-based intelligent chest card collection and analysis method and system. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, it has been widely applied in the service field. In the service process, the service quality of service personnel is inspected, and the audio in the service process is collected through intelligent chest cards, and intelligent analysis is performed combined with AI technology, which can greatly improve the service quality.
[0003] The existing AI inspection system mainly relies on the keyword matching mode to realize the judgment and inspection process of service quality. The semantic understanding is shallow, only the preset keywords and the preset process can be identified, and the context semantics cannot be understood. It is powerless to complex semantics and implicit information such as irony and implication, resulting in a large number of missed inspections and false positives. Only relying on a single keyword or part of the semantic ability of a small model, the detection dimension is relatively single, and other abnormal situations cannot be identified. The customer's emotions, intentions and other information cannot be identified, and the service quality cannot be evaluated comprehensively. In addition, the traditional inspection only scans the fragments, cannot fully understand the dialogue content, resulting in missed inspections and false positives. The overall structure and process of the dialogue cannot be grasped, the problems cannot be accurately positioned, the traditional inspection relies on manual maintenance rules, the iteration cost is high, the cycle is long, the rules are not updated in time, resulting in inaccurate inspection results, affecting the service quality. Moreover, the sound collecting device of the traditional inspection system is matched with the telephone, and in a noisy environment, the sound collecting effect cannot be guaranteed. Therefore, the quality of the audio source cannot be well guaranteed, thereby affecting the audio transcription text, and finally affecting the model matching result, which affects the accuracy of the inspection. Therefore, how to improve the quality of the intelligent chest card sound collection and improve the accuracy of the service quality judgment of the inspection system is the fundamental problem to be solved by the application. SUMMARY
[0004] In order to improve the quality of the intelligent chest card sound collection and improve the accuracy of the service quality judgment of the inspection system, the application provides an AI-based intelligent chest card collection and analysis method and system.
[0005] In the first aspect, the application provides an AI-based intelligent chest card collection and analysis method, which adopts the following technical scheme:
[0006] The AI-based intelligent chest card collection and analysis method comprises:
[0007] The communication between the device management service module and the chest card device is realized through the emqx service module, and the information configuration of the chest card device is realized through the device management service module;
[0008] The sound information is collected by the chest card device;
[0009] The sound information is uploaded to a quality inspection analysis platform, the sound information is processed through the quality inspection analysis platform, audio transcription text is obtained, the audio transcription text is analyzed based on an AI large model, and a service judgment result is obtained.
[0010] By using the above technical solution, AI analysis is performed through the audio transcription text, which can improve the understanding of semantics and hidden information compared with the traditional keyword matching method, and can improve the accuracy of service quality judgment.
[0011] Optionally, the chest card device comprises a microphone array and a structured light module, and the process of collecting sound information by the chest card device comprises:
[0012] The position information and the face information of the person are recognized by the structured light module, and it is judged whether the person is in a speaking state according to the face information of the person;
[0013] The real-time noise intensity is obtained based on a preset noise recognition strategy;
[0014] The sensitivity of the microphone array is dynamically adjusted according to the position information of the person in the speaking state and the real-time noise intensity, and sound information is collected.
[0015] By using the above technical solution, the intelligent chest card recognizes the position information and the face information of the person through the structured light module, and then dynamically adjusts the sensitivity of the microphone array according to the position information of the speaking person and the real-time noise intensity, so that the intelligibility of human voice collection can be greatly improved compared with the microphone sound collection method with fixed sensitivity, and the accuracy of service quality judgment can be improved.
[0016] Optionally, the process of dynamically adjusting the sensitivity of the microphone array comprises:
[0017] Through the model:
[0018]
[0019] The sensitivity Sens(t) of the microphone array is calculated and obtained;
[0020] Wherein d(t) is the real-time distance between the person in the speaking state and the chest card device, is a reference distance, =1m, is the sound pressure level at the reference distance is the sound pressure level at the reference distance is the human ear threshold, , is the environmental real-time noise sound pressure level, is a position sensitivity reference function, >0, a noise sensitivity reference function, a corresponding minimum level.
[0021] By adopting the technical scheme, the real-time microphone array sensitivity can be obtained through the acquired personnel distance information and real-time noise information, and the quality of the sound collection can be improved by collecting sound according to the microphone array sensitivity, thereby improving the accuracy of the service quality judgment.
[0022] Optionally, the preset noise recognition strategy comprises:
[0023] obtaining a noise intensity change curve of each group of microphones in the microphone array through a noise recognition algorithm;
[0024] calculating a correlation coefficient of noise intensity of any two groups of microphones, and obtaining a noise group microphone according to the correlation coefficient of noise intensity of each group of microphones and other groups of microphones and the noise intensity change curve of each group of microphones;
[0025] turning off the noise group microphone, and obtaining an environmental real-time noise sound pressure level according to sound data of the remaining groups of microphones.
[0026] By adopting the technical scheme, the noise group microphone is obtained according to the correlation coefficient of noise intensity of each group of microphones and other groups of microphones and the noise intensity change curve of each group of microphones, and the noise group microphone is turned off, so that the influence of noise on the quality of sound collection can be greatly reduced, that is, the quality of sound collection is improved.
[0027] Optionally, the process of obtaining the noise group microphone comprises:
[0028] through a model:
[0029]
[0030] calculating a noise direction real-time matching value of the ith group of microphones , and selecting a microphone group corresponding to a maximum noise direction real-time matching value as the noise group microphone; wherein,
[0031] is a preset fixed time period, is a noise intensity change curve of the ith group of microphones, is an average of all noise intensity change curves, is an adjustment coefficient, m is the number of microphone groups, is a sum of correlation coefficients of the ith group of microphones and other groups of microphones.
[0032] By adopting the technical scheme, when the noise direction real-time matching value is larger, it indicates that the noise intensity of the i-th group of microphones is stronger, and the correlation of the remaining groups of microphones is higher, so that the group of microphones corresponding to the maximum noise direction real-time matching value is selected as the noise group microphone and is closed, which can greatly reduce the influence of noise on the sound quality, and further improve the quality of sound collection.
[0033] Optionally, the calculation process of the correlation coefficient comprises:
[0034] acquiring the noise intensity of the y-th group of microphones at the x-th time point and the noise intensity of the z-th group of microphones at the x-th time point according to a fixed time interval q time points in a period;
[0035] by the model:
[0036]
[0037] obtaining the correlation coefficient of the y-th group of microphones and the z-th group of microphones by calculation .
[0038] wherein x [1, q], is the x-th time point, is the noise intensity of the y-th group of microphones at the x-th time point, is the average noise intensity of the y-th group of microphones at the q time points, is the noise intensity of the z-th group of microphones at the x-th time point, is the average noise intensity of the z-th group of microphones at the q time points.
[0039] By adopting the technical scheme, the judgment of the noise group microphone is realized through the size of the correlation coefficient.
[0040] Optionally, the process of obtaining the audio transcription text comprises:
[0041] periodically scanning the sound information in the path, and performing voiceprint recognition on the sound information through a voiceprint service module;
[0042] performing role separation through a role separation service module, and obtaining the audio transcription text of the separated roles through an ASR speech recognition transcription service.
[0043] By adopting the technical scheme, the role separation process can be realized, and the accuracy of the judgment result is improved when analyzed by the AI large model.
[0044] Optionally, the process of analyzing the audio transcription text based on the AI large model comprises:
[0045] S1, configuring the AI large model, comprising:
[0046] The AI generation configuration module configures analysis preset information, including an AI type and a prompt word;
[0047] An AI quality inspection rule is configured through an AI quality inspection rule module.
[0048] S2, the AI large model configured is used to analyze the audio transcribed text.
[0049] By adopting the above technical solution, accurate evaluation of services can be realized.
[0050] In a second aspect, the application provides an AI-based intelligent name badge collection and analysis system, which adopts the following technical solution:
[0051] The AI-based intelligent name badge collection and analysis system comprises an emqx service module, a device management service module, a name badge device, a quality inspection analysis platform, and an AI large model.
[0052] The device management service module communicates with the name badge device through the emqx service module.
[0053] The device management service module is configured to configure information for the name badge device.
[0054] The name badge device is configured to collect sound information.
[0055] The quality inspection analysis platform is configured to process the uploaded sound information to obtain an audio transcribed text.
[0056] The AI large model is configured to analyze the audio transcribed text to obtain a service judgment result.
[0057] In summary, the application has at least one of the following beneficial technical effects:
[0058] The application analyzes AI through the audio transcribed text, which can improve the understanding of semantics and hidden information and the accuracy of service quality judgment compared with the traditional keyword matching method. In addition, the intelligent name badge identifies personnel position information and personnel face information through a structured light module, and then dynamically adjusts the sensitivity of the microphone array according to the position information of the speaker and the real-time noise intensity, so that the intelligibility of human voice collection can be greatly improved compared with the microphone pickup method with fixed sensitivity, thereby improving the accuracy of service quality judgment. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 is a flowchart of the AI-based intelligent name badge collection and analysis method in the application.
[0060] Figure 2 is a flowchart of the AI-based intelligent name badge collection and analysis system in the application. DETAILED DESCRIPTION
[0061] Embodiments of the present application are described in detail below with reference to the attached drawing figures, wherein examples of embodiments are illustrated.
[0062] In the description of the present specification, the description referring to the terms "certain embodiments", "one embodiment", "some embodiments", "illustrative embodiments", "example", "specific example" or "some examples" means that the particular feature, structure, material or characteristic being described is included in at least one embodiment or example of the present application. The illustrative descriptions of such terms in this specification are not necessarily referring to the same embodiment or example. Furthermore, the described particular features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0063] The embodiment of the present application discloses an AI-based intelligent chest card collection and analysis method, referring to Figure 1 , comprising: realizing the communication between the device management service module and the chest card device through the emqx service module, and configuring information for the chest card device through the device management service module; collecting sound information through the chest card device; uploading the sound information to the quality inspection analysis platform, processing the sound information through the quality inspection analysis platform, obtaining audio transcription text, analyzing the audio transcription text based on AI large model, and obtaining service judgment result. Wherein, the chest card device comprises a microphone array and a structured light module, and the process of collecting sound information through the chest card device comprises: identifying personnel position information and personnel face information through the structured light module, judging whether it is in a speaking state according to the personnel face information; obtaining real-time noise intensity based on a preset noise recognition strategy; dynamically adjusting the sensitivity of the microphone array according to the position information of the personnel in the speaking state and the real-time noise intensity, and obtaining the collected sound information. Through the above scheme, it can be known that the intelligent chest card in the embodiment identifies personnel position information and personnel face information through the structured light module, and then dynamically adjusts the sensitivity of the microphone array according to the position information of the speaking personnel and the real-time noise intensity. Therefore, compared with the microphone receiving mode with fixed sensitivity, the intelligibility of human voice collection can be greatly improved, and the accuracy of service quality judgment can be improved. In addition, the AI analysis is performed in the form of audio transcription text in the embodiment, and the AI analysis can be implemented by using the existing AI large model, such as GPT system large model. Compared with the traditional keyword matching mode, the understanding ability of semantics and hidden information can be improved, and the accuracy of service quality judgment can be improved.
[0064] It should be noted that the intelligent chest card in the embodiment will request the consent and authorization of the customer before collecting the face information and sound information, and the collected face information and sound information are only used for service quality judgment and improvement, and are not used for other purposes.
[0065] wherein the process of dynamically adjusting the sensitivity of the microphone array comprises: obtaining the sensitivity Sens(t) of the microphone array by a model:
[0066]
[0067] obtaining the sensitivity Sens(t) of the microphone array by a model: is a reference distance, = 1 m, is a reference distance of the sound pressure level at the position, is the human ear threshold, , is the environmental real-time noise sound pressure level, is a position sensitivity reference function, > 0, is a noise sensitivity reference function, , is corresponding minimum level, the position sensitivity reference function and the noise sensitivity reference function are obtained according to test data, the position sensitivity reference function is obtained according to the best sound pickup sensitivity under different distance states in the test data, and the noise sensitivity reference function is obtained according to the sound pickup sensitivity adjustment value corresponding to different noise levels in the test data. Therefore, by obtaining the distance information of the person and the real-time noise information, the real-time sensitivity of the microphone array can be obtained, and the sound pickup quality can be improved by sound pickup according to the sensitivity of the microphone array, thereby improving the accuracy of the service quality judgment.
[0068] In addition, in one embodiment, a preset noise recognition strategy is provided, which comprises: obtaining the noise intensity change curve of each group of microphones in the microphone array by a noise recognition algorithm; calculating the correlation coefficient of the noise intensity of any two groups of microphones, obtaining the noise group microphone according to the correlation coefficient of the noise intensity of each group of microphones and other groups of microphones and the noise intensity change curve thereof; closing the noise group microphone, and obtaining the environmental real-time noise sound pressure level according to the sound data of the remaining groups of microphones. Since each group of microphones in the microphone array picks up sound in different directions (there is overlap in microphone pickup), by the above method, the noise group microphone is obtained according to the correlation coefficient of the noise intensity of each group of microphones and other groups of microphones and the noise intensity change curve thereof, and is closed, so that the influence of noise on the sound pickup quality can be greatly reduced, that is, the sound pickup quality is improved.
[0069] It should be noted that the sensitivity Sens(t) of the microphone array is obtained for the part of the microphone array that is turned on, and the noise group microphone that is turned off does not need to be adjusted. In addition, in the case where the speaker position and the noise source are in the same direction, the sound can still be received through the non-noise group microphone, and the sound quality is still stronger than that of the noise group microphone that is turned on.
[0070] wherein the process of obtaining the noise group microphone includes:
[0071]
[0072] The noise direction real-time matching value of the ith group of microphones is calculated , wherein is a preset fixed period, which is selected according to experience, and the unit is seconds, is the noise intensity change curve of the ith group of microphones, is the average of all group noise intensity change curves, is an adjustment coefficient, which is determined according to the test data and the multiple of the numerical range and the weight, and m is the number of microphone groups, is the sum of the correlation coefficients of the ith group of microphones and other groups of microphones, so when the noise direction real-time matching value is larger, it means that the noise intensity of the ith group of microphones is stronger, and the correlation of the remaining groups of microphones is higher (the noise data obtained by the noise group microphone is the clearest, so the correlation between the remaining groups and the noise group microphone will be higher than the correlation between the remaining groups), so the microphone group corresponding to the maximum noise direction real-time matching value is selected as the noise group microphone and is turned off, which can greatly reduce the influence of noise on the sound quality, and further improve the sound quality.
[0073] In addition, the calculation process of the correlation coefficient includes: obtaining at q time points in a fixed time interval; through the model:
[0074]
[0075] The correlation coefficient between the yth group of microphones and the zth group of microphones is calculated ; wherein x∈[1, q], is the xth time point, is the noise intensity of the yth group of microphones at the xth time point, is the average noise intensity of the yth group of microphones at q time points, is the noise intensity of the zth group of microphones at the xth time point, Let be the mean noise intensity of the z-th microphone at q time points, therefore the correlation coefficient is... The closer the value is to 1, the stronger the correlation. Therefore, the magnitude of the correlation coefficient can be used to determine the noise group microphone.
[0076] It should be noted that the correlation coefficient This represents the correlation between the y-th microphone group and the z-th microphone group. When y represents the ith microphone group and z represents the remaining m-1 microphone groups, m-1 correlation coefficients can be obtained. The sum of the m-1 correlation coefficients is the result. .
[0077] In one embodiment, the process of obtaining audio-to-text transcription includes: periodically scanning sound information in the path, performing voiceprint recognition on the sound information through a voiceprint service module; performing role separation through a role separation service module; and obtaining the audio-to-text transcription with separated roles through an ASR speech recognition transcription service. Through the above process, role separation can be achieved, thereby improving the accuracy of the judgment results when analyzed using an AI large-scale model. Furthermore, the process of analyzing the audio-to-text transcription based on the AI large-scale model includes: S1, configuring the AI large-scale model, including configuring preset analysis information through an AI generation configuration module. The preset information includes AI type and prompt words, etc. S1. Configure AI quality inspection rules through the AI quality inspection rule module; S2. Analyze the audio-to-text transcript using the configured AI big model. The AI big model can accurately identify complex semantics and implicit information through contextual reasoning and deep semantic understanding, greatly improving the accuracy of the analysis; Finally, the analysis results are given to the business optimization module, which optimizes the data based on the results. Because the big model can continuously learn from new data and adapt to new business types and quality standards, the model evolves autonomously with low iteration costs and short cycles. It can adapt to business changes in a timely manner, improve quality inspection results, reduce iteration costs, and greatly improve the work efficiency of operations personnel and reduce operating costs.
[0078] This application also discloses an AI-based intelligent name tag collection and analysis system. Please refer to the appendix. Figure 2 As shown, it includes an EMQX service module, an equipment management service module, name badge devices, a quality inspection and analysis platform, and an AI big data model. The equipment management service module communicates with the name badge devices through the EMQX service module. The equipment management service module is used to configure information on the name badge devices. The name badge devices are used to collect audio information. The quality inspection and analysis platform is used to process the uploaded audio information to obtain audio-to-text transcription. The AI big data model is used to analyze the audio-to-text transcription to obtain business judgment results.
[0079] Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary, and are not to be interpreted as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. An AI-based intelligent chest badge collection analysis method, characterized in that, The system adopts the AI-based intelligent chest card collection and analysis method of any one of claims 1-5, comprising an emqx service module, a device management service module, a chest card device, a quality inspection analysis platform, and an AI large model. The device management service module communicates with the chest card device through the emqx service module. The device management service module is configured to configure information for the chest card device. The chest card device is configured to collect sound information. The quality inspection analysis platform is configured to process the uploaded sound information to obtain audio transcription text. The AI large model is configured to analyze the audio transcription text to obtain a business judgment result. The chest card device comprises a microphone array and a structured light module. The process of collecting sound information by the chest card device comprises: The structured light module identifies personnel position information and personnel facial information, and determines whether the personnel are in a speaking state based on the personnel facial information. The sensitivity of the microphone array is dynamically adjusted based on the real-time noise intensity and the position information of the personnel in the speaking state to obtain the collected sound information. ; The process of dynamically adjusting the sensitivity of the microphone array comprises: Wherein, d(t) is the real-time distance between the speaker and the chest badge device, is the reference distance, = 1m, is the reference distance of the sound pressure level at the position, is the human ear hearing threshold, , is the environmental real-time noise sound pressure level, is the position sensitivity reference function, > 0, is the noise sensitivity reference function, , is the corresponding minimum level; The model is used to calculate the sensitivity Sens(t) of the microphone array. The preset noise recognition strategy comprises: A noise recognition algorithm is used to obtain the noise intensity change curve of each group of microphones in the microphone array. The correlation coefficient of the noise intensity of any two groups of microphones is calculated, and the noise group microphone is obtained based on the correlation coefficient of the noise intensity of each group of microphones and other groups of microphones and the noise intensity change curve thereof. 2.The AI-based intelligent chest plate acquisition analysis method of claim 1, wherein, The noise group microphone is turned off, and the real-time noise sound pressure level of the environment is obtained based on the sound data of the remaining groups of microphones. The process of obtaining the noise group microphone comprises: ; Calculating the real-time matching value of the noise direction of the ith group of microphones , selecting the real-time matching value of the noise direction The group of microphones corresponding to the maximum value is the group of microphones of the noise wherein, is a preset fixed period, is the noise intensity variation curve of the i-th group of microphones, is the average of the noise intensity variation curves of all groups, is an adjustment coefficient, m is the number of groups of microphones, is the sum of the correlation coefficients of the i-th group of microphones with other groups of microphones. 3.The AI-based intelligent chest plate acquisition analysis method of claim 2, wherein, The model is used. acquired at fixed time intervals q time points within a period The process of calculating the correlation coefficient comprises: ; computing a correlation coefficient of the yth group of microphones and the zth group of microphones ; wherein x e [1, q], is the xth time point, is the noise intensity of the yth group of microphones at the xth time point, is the average noise intensity of the yth group of microphones at the q time points, is the noise intensity of the zth group of microphones at the xth time point, is the average noise intensity of the zth group of microphones at the q time points. 4.The AI-based intelligent chest plate acquisition analysis method of claim 1, wherein, The model is used. The process of obtaining the audio transcription text comprises: Voiceprint recognition is performed on the sound information by a voiceprint service module. 5.The AI-based intelligent chest plate acquisition analysis method of claim 1, wherein, Role separation is performed by a role separation service module, and audio transcription text of the separated roles is obtained by an ASR speech recognition transcription service. The process of analyzing the audio transcription text based on the AI large model comprises: S1, configuring the AI large model, comprising: Analysis preset information is configured by an AI generation configuration module, including AI type and prompt words. AI quality inspection rules are configured by an AI quality inspection rule module.
6. An AI-based intelligent chest badge collection analysis system, characterized by, S2, analyzing the audio transcription text by the configured AI large model. The system adopts the AI-based intelligent chest card collection and analysis method of any one of claims 1-5, comprising an emqx service module, a device management service module, a chest card device, a quality inspection analysis platform, and an AI large model; The device management service module communicates with the chest card device through the emqx service module. The device management service module is configured to configure information for the chest card device. The chest card device is configured to collect sound information. The quality inspection analysis platform is configured to process the uploaded sound information to obtain audio transcription text. The AI large model is configured to analyze the audio transcription text to obtain a business judgment result.
Citation Information
Patent Citations
Pickup control method, device and system, equipment and medium
CN112053701A
Processing method and system based on large model and work card equipment verbal skill quality inspection
CN118248164A