Automobile parts size ai measurement and defect detection integrated platform
By integrating voice interaction, hardware acquisition, and feature fusion into a unified inspection platform, the problems of low efficiency and system fragmentation in traditional inspection methods have been solved, enabling efficient and accurate measurement of automotive parts dimensions and defect detection, and adapting to complex industrial environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AOLIQI TECH CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-12
AI Technical Summary
Traditional automotive parts inspection methods are inefficient, lack precision, and are easily affected by subjective factors. The dimensional measurement and defect detection systems are disconnected, making it impossible to achieve high-precision and high-efficiency quality control.
An integrated platform comprising a voice interaction module, a hardware acquisition module, a feature fusion module, and an integrated detection module is adopted to achieve deep integration of full-process voice interaction, dimensional measurement, and defect detection. Through simultaneous data acquisition via multispectral imaging and 3D scanning, combined with deep learning networks to extract joint feature vectors, integrated detection is performed.
It enables efficient and accurate dimensional measurement and defect detection in complex industrial environments, improves operational convenience and inspection efficiency, and ensures comprehensive and accurate inspection of automotive parts.
Smart Images

Figure CN122192159A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial automation inspection technology, specifically to an integrated platform for AI measurement and defect detection of automotive parts dimensions. Background Technology
[0002] In the automotive parts manufacturing process, dimensional measurement and defect detection are core aspects of ensuring product quality. Traditional inspection methods mainly rely on manual visual inspection or mechanical measuring tools, which suffer from low efficiency, insufficient accuracy, and susceptibility to subjective factors. With the development of industrial automation, existing inspection platforms generally use handles or touch screens for operation and control. However, in actual production environments, operators need to wear protective gloves, and the work area is often contaminated with oil, leading to sluggish or even malfunctioning touch operation responses, seriously affecting the continuity of the inspection process. At the same time, the high background noise environment in industrial sites makes it difficult for voice recognition systems to work stably. Operators cannot complete operations such as starting inspection tasks, setting parameters, and querying results through natural voice commands, and are forced to rely on physical interaction devices.
[0003] Furthermore, dimensional measurement systems and defect detection systems typically operate as independent modules, with data acquisition and processing flows fragmented. This prevents joint analysis of the two-dimensional texture and three-dimensional geometric features of the same component, resulting in a lack of correlation and verification between critical dimensional parameter calculations and surface defect identification. Consequently, contradictory detection results or the omission of complex defects are prone to occur. This separate architecture not only increases equipment deployment costs but also reduces overall inspection efficiency, making it difficult to meet the high-precision and high-efficiency quality control requirements of modern automobile manufacturing.
[0004] Therefore, there is an urgent need to develop an integrated inspection platform that can adapt to complex industrial environments, support full-process voice interaction, and achieve deep integration of dimensional measurement and defect detection functions. Summary of the Invention
[0005] In view of this, the present invention provides an integrated AI measurement and defect detection platform for automotive parts dimensions, which has the advantages of being able to adapt to complex industrial environments, supporting full-process voice interaction, and achieving deep integration of dimension measurement and defect detection functions, thereby improving detection efficiency and accuracy.
[0006] This invention provides an integrated platform for AI measurement and defect detection of automotive parts dimensions, comprising: The voice interaction module is used to receive and recognize the operator's voice commands, convert the voice commands into control commands, and feed back the detection results to the operator through voice synthesis. The hardware acquisition module, connected to the voice interaction module, responds to the acquisition start command in the control command and includes a multispectral imaging unit, a high-precision three-dimensional scanning unit and a flexible transmission unit. The multispectral imaging unit and the high-precision three-dimensional scanning unit are arranged sequentially or collaboratively along the transmission direction of the flexible transmission unit, and are used to acquire multispectral two-dimensional image data and three-dimensional point cloud data of automotive parts synchronously or in time-division. The feature fusion module, connected to the hardware acquisition module, is used to align the multispectral two-dimensional image data and the three-dimensional point cloud data of the same automotive part in terms of spatial coordinates, and to extract a joint feature vector that fuses two-dimensional texture features and three-dimensional geometric features based on a deep learning network. An integrated detection module, connected to the feature fusion module, has a built-in shared joint feature vector processing channel for parallel or pipelined execution. The dimensional measurement task, based on the three-dimensional geometric feature information in the joint feature vector, calculates the key dimensional parameters of the parts through sub-pixel level edge detection and geometric fitting algorithms; The defect detection task, based on the two-dimensional texture features and three-dimensional geometric features in the joint feature vector, identifies texture defects and three-dimensional structural defects on the surface of parts through a pre-trained defect classification and segmentation model. The decision control module, connected to the integrated detection module and the voice interaction module, is used to generate detection result data based on the combined results of the size measurement task and the defect detection task, and send the detection result data to the voice interaction module for voice feedback. At the same time, it generates sorting control instructions for the parts on the flexible conveying unit based on the detection result data.
[0007] As can be seen from the above, the AI-integrated platform for automotive component size measurement and defect detection provided in this application integrates a voice interaction module to receive and convert voice commands, a hardware acquisition module to acquire multispectral images and 3D point cloud data, a feature fusion module to align and extract joint feature vectors, an integrated detection module to execute size measurement and defect detection tasks in parallel, and a decision control module to generate detection results and sorting instructions. This achieves integrated processing of full-process voice interaction and size measurement and defect detection, and has the advantages of being able to adapt to complex industrial environments, supporting full-process voice interaction, and achieving deep integration of size measurement and defect detection functions. Attached Figure Description
[0008] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0009] Figure 1 This is a structural framework diagram of the integrated AI measurement and defect detection platform for automotive parts dimensions according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the voice interaction module in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] Traditional AI inspection platforms for automotive parts typically use handles or touchscreens for operation, which are cumbersome and susceptible to interference from gloves and oil, affecting ease of operation and inspection efficiency.
[0012] In this regard, such as Figure 1 As shown, this application proposes an integrated platform for AI measurement and defect detection of automotive parts dimensions, including: The voice interaction module is used to receive and recognize the operator's voice commands, convert the voice commands into control commands, and feed back the detection results to the operator through voice synthesis.
[0013] The hardware acquisition module, connected to the voice interaction module, responds to the acquisition start command in the control instructions. It includes a multispectral imaging unit, a high-precision 3D scanning unit, and a flexible transmission unit. The multispectral imaging unit and the high-precision 3D scanning unit are arranged sequentially or collaboratively along the transmission direction of the flexible transmission unit, and are used to acquire multispectral 2D image data and 3D point cloud data of automotive parts synchronously or in time-division.
[0014] The feature fusion module, connected to the hardware acquisition module, is used to align the spatial coordinates of multispectral two-dimensional image data and three-dimensional point cloud data of the same automotive part, and extract a joint feature vector that fuses two-dimensional texture features and three-dimensional geometric features based on a deep learning network.
[0015] The integrated detection module, connected to the feature fusion module, has a built-in shared joint feature vector processing channel for parallel or pipelined execution. The dimensional measurement task, based on the three-dimensional geometric feature information in the joint feature vector, calculates the key dimensional parameters of the parts through sub-pixel level edge detection and geometric fitting algorithms; The defect detection task, based on the two-dimensional texture features and three-dimensional geometric features in the joint feature vector, identifies texture defects and three-dimensional structural defects on the surface of parts through a pre-trained defect classification and segmentation model.
[0016] The decision control module, connected to the integrated detection module and the voice interaction module, is used to generate detection result data based on the combined results of the dimensional measurement task and the defect detection task, and send the detection result data to the voice interaction module for voice feedback. At the same time, it generates sorting control instructions for the parts on the flexible conveyor unit based on the detection result data.
[0017] The voice interaction module is the interface for human-computer interaction. Its function is to receive the operator's voice input, convert it into machine-recognizable instructions, and provide feedback to the operator in voice form.
[0018] The hardware acquisition module is responsible for acquiring the raw data of the automotive parts to be tested. It contains units for different data acquisition methods and can start the data acquisition process according to control commands.
[0019] Multispectral two-dimensional image data refers to a collection of images acquired by a multispectral imaging unit that reflect the two-dimensional texture information of the surface of a component under different spectral bands.
[0020] Three-dimensional point cloud data refers to a set of a large number of points with three-dimensional coordinates obtained by a high-precision three-dimensional scanning unit, used to describe the geometry and spatial structure of parts.
[0021] The joint feature vector is a high-dimensional data representation generated after processing by the feature fusion module. It integrates the two-dimensional texture information and three-dimensional geometric information of the parts, providing a comprehensive data foundation for subsequent detection tasks.
[0022] The integrated inspection module is the core of the specific inspection task. It is designed to handle two types of tasks, namely dimensional measurement and defect detection, simultaneously or alternately, and to share a unified feature processing channel.
[0023] The decision control module is responsible for making comprehensive judgments based on the test results, generating corresponding control instructions (such as sorting instructions) and feedback information to achieve control and management of the entire testing process.
[0024] Specifically, the voice interaction module is the core of the human-machine interface. It receives operator voice commands, converts them into machine-executable control commands through voice recognition and understanding, and is responsible for broadcasting the detection results to the operator in voice form, achieving closed-loop voice interaction. The hardware acquisition module starts upon receiving the voice command. Multispectral imaging units (such as visible light and infrared cameras) and high-precision 3D scanning units (such as structured light scanners) are sequentially or collaboratively arranged along a flexible conveyor unit (such as a conveyor belt) to acquire multispectral 2D images and 3D point cloud data of the same component, providing multimodal input for subsequent analysis. The feature fusion module aligns the spatial coordinates of these two heterogeneous data sources (e.g., projecting the 3D point cloud onto a 2D image plane using calibration parameters) and extracts joint feature vectors using a deep learning network (such as a two-stream convolutional neural network). F The joint feature vector F It simultaneously encodes two-dimensional texture information (color, edges, texture) and three-dimensional geometric information (depth, curvature, normal). The integrated detection module shares this joint feature vector. F Two tasks can be executed in parallel or pipelined manner: the dimensional measurement task utilizes joint feature vectors. F The geometric features in the image are used to calculate key dimensions through sub-pixel edge detection (such as Canny+ interpolation) and geometric fitting (such as least squares circle fitting); the defect detection task utilizes joint feature vectors. F The system identifies all features in the data using pre-trained defect classification networks (such as ResNet) and segmentation networks (such as U-Net) to identify surface defects (scratches, cracks, burrs, etc.). The decision control module integrates size and defect results to generate a quality judgment. On one hand, it sends the result text to the voice interaction module for voice feedback; on the other hand, it generates sorting instructions to control the transmission unit to automatically sort the parts. Through this integrated design, the platform achieves fully intelligent operation from voice command input and automatic detection to voice result feedback.
[0025] This embodiment provides an integrated AI-powered platform for automotive component dimensional measurement and defect detection. By introducing voice interaction, it effectively solves the problems of cumbersome operation and susceptibility to interference from gloves and oil contamination found in traditional inspection platforms. Operators can conveniently control the inspection process and obtain result feedback via voice commands, significantly improving operational ease of use in industrial environments. Simultaneously, multimodal data acquisition and fusion, combined with integrated dimensional measurement and defect detection, ensures comprehensive and accurate inspection of automotive components.
[0026] In one alternative implementation, such as Figure 2 As shown, the voice interaction module includes: a voice acquisition unit, a preprocessing unit, a voice recognition unit, a natural language understanding unit, a dialogue management unit, and a voice synthesis unit.
[0027] The voice acquisition unit is used to capture the raw voice signal emitted by the operator. It can employ a high-sensitivity microphone array, such as an array of multiple omnidirectional or directional microphones, to ensure effective speech reception from different directions and distances. The speech acquisition unit typically has wide frequency response characteristics, enabling it to acquire high-quality raw audio data containing rich speech information.
[0028] The preprocessing unit is connected to the voice acquisition unit, and its main function is to process the raw voice signal. Optimization is performed to improve the accuracy of subsequent speech recognition. This includes, but is not limited to, operations such as echo cancellation, noise suppression, beamforming, and speech activity detection on the speech signal. Specifically, the preprocessing unit... Processing is performed to improve the signal-to-noise ratio: echo cancellation (AEC) eliminates interference signals played from the speaker; noise suppression (NS) reduces ambient noise; beamforming (BF) enhances speech in the target direction; and voice activity detection (VAD) determines the start and end points of speech and extracts clean speech segments. .
[0029] The speech recognition unit is connected to the preprocessing unit and is responsible for converting the preprocessed enhanced speech segments into text instructions that can be understood by the machine. T Speech recognition units are typically based on acoustic and language models, mapping acoustic features in speech signals to corresponding text sequences. Besides end-to-end models, hybrid models combining Hidden Markov Models (HMMs) and Deep Neural Networks (DNNs), or sequence-to-sequence models based on Recurrent Neural Networks (RNNs) or Long Short-Term Memory Networks (LSTMs) can also be used to achieve speech-to-text conversion.
[0030] The natural language understanding unit is connected to the speech recognition unit, and its function is to process the text instructions output by the speech recognition unit. T Deep semantic analysis is performed. The natural language understanding unit can identify key information in the text, such as the operator's intent (e.g., "start detection", "stop detection") and related entities (e.g., "part type", "measurement item", "value").
[0031] The dialogue management unit is connected to the natural language understanding unit and is responsible for maintaining the context and state of the entire dialogue. Based on the intent and slot information extracted by the natural language understanding unit, it determines the current progress of the dialogue and decides on the next action. For example, when key information is detected as missing, the dialogue management unit generates follow-up questions. Q It guides the operator to provide necessary information; once all necessary information has been obtained, it generates corresponding control instructions. C .
[0032] The speech synthesis unit is connected to the dialogue management unit and the decision control module, and is used to convert text information (such as detection result data or follow-up text) into natural and fluent speech signals and play them to the operator.
[0033] Through the above technical solution, this application provides a well-structured and powerful voice interaction module, effectively solving the problems of inaccurate voice command recognition and low human-computer interaction efficiency in industrial environments. The voice acquisition unit can capture the operator's original voice with high fidelity, while the preprocessing unit can effectively suppress environmental noise and echo, significantly improving the clarity and signal-to-noise ratio of the voice signal, laying a solid foundation for subsequent processing. Based on this, the voice recognition unit can accurately convert the enhanced voice segments into text commands, overcoming the impact of complex acoustic environments on recognition accuracy. The natural language understanding unit further performs semantic analysis on the text commands, accurately identifying the operational intent and key parameters, enabling the system to understand the operator's true needs, rather than just literal words. The dialogue management unit maintains the dialogue state, intelligently guiding the operator to complete command input and initiating follow-up questions when necessary, ensuring the completeness of all necessary information, thereby improving the efficiency and accuracy of command generation. Finally, the voice synthesis unit can clearly and naturally feed back the detection results or follow-up text to the operator, achieving a smooth and efficient human-computer voice interaction closed loop. This sophisticated voice interaction mechanism enables operators to interact with the integrated AI measurement and defect detection platform for automotive parts dimensions in a more natural and convenient way. Especially in industrial scenarios where hands are inconvenient to operate or vision is limited, it greatly improves the ease of operation, work efficiency, and overall system intelligence, ensuring that the platform can stably and reliably execute various measurement and inspection tasks.
[0034] In one alternative implementation, the preprocessing unit is configured as follows: Echo cancellation is performed using a normalized minimum mean square adaptive filter; Noise suppression is achieved using spectral subtraction based on logarithmic minimum mean square error; A time difference of arrival estimation algorithm based on generalized cross-correlation-phase transform is used to locate the sound source, and a multi-channel signal is weighted and summed based on a delay summation beamformer, with the beam pointing in the direction of the sound source.
[0035] Specifically, to effectively eliminate echoes in the speech signal, the preprocessing unit is configured to use a Normalized Least Mean Square (NLMS) adaptive filter for echo cancellation. The NLMS adaptive filter estimates and cancels echo paths by adaptively adjusting its coefficients, thereby separating clean near-end speech. Its coefficient update formula is as follows: ; in, This is the filter coefficient vector; This is the updated coefficient vector of the filter; This is the step size factor (ranging from 0.05 to 0.15), used to balance the update speed and stability of the filter coefficients; Reference signal (speaker output); The error signal is the microphone signal minus the estimated echo. To prevent division by zero for small constants; filter length L The echo path modeling capability was determined and set within the range of 1024 to 2048 to cover echo paths in different environments.
[0036] To suppress background noise prevalent in industrial environments, the preprocessing unit also employs noise suppression based on the logarithmic minimum mean square error (LogMMSE) spectral subtraction method. This method estimates the noise power spectrum. Updated via recursive averaging: ; in, For frame index, For frequency points, For noisy speech spectra, update factors (0.90~0.98) control the tracking speed of noise estimation.
[0037] Furthermore, to address situations where the operator's position is not fixed or there is interference from multiple sound sources, the preprocessing unit is also equipped with a sound source localization function. This function uses a time difference of arrival estimation algorithm based on generalized cross-correlation-phase transform (GCC-PHAT) to accurately calculate the position of the sound source relative to the microphone array. After obtaining the sound source direction, the preprocessing unit further performs weighted summation of the multi-channel signals based on a delay-summing beamformer and points the beam towards the sound source direction. This beamformer, by appropriately delaying and weighting the signals from different microphones, ensures that the signals from the target sound source direction are coherently superimposed, while noise and interference signals from other directions are effectively suppressed.
[0038] Among them, the time difference of arrival The calculation expression is: ; in, , The Fourier transform of the signals from the two microphones. It is the angular frequency. According to... Calculate and sum the delays of each channel to enhance the speech in the target direction.
[0039] Through the aforementioned technical solutions, the preprocessing unit provides an efficient solution to the complex acoustic challenges in industrial environments. The normalized minimum mean square adaptive filter accurately eliminates echoes, ensuring that operator commands are not confused by their own speech reflections. Spectral subtraction based on logarithmic minimum mean square error effectively filters out background noise, ensuring the speech signal remains clear even in noisy environments. More importantly, the combination of source localization achieved through generalized cross-correlation-phase transformation and a delay-summing beamformer enables the system to intelligently focus on the operator's speech. Even if the operator's position changes, it significantly improves the signal-to-noise ratio of the target speech, thus providing high-quality input for the subsequent speech recognition unit. This greatly improves the accuracy and robustness of voice command recognition, reduces the risk of misoperation, ensures the stability and reliability of the platform in practical industrial applications, and ultimately improves overall operational efficiency and user experience.
[0040] In one alternative implementation, the speech recognition unit employs an end-to-end speech recognition model based on the Conformer-CTC architecture. This speech recognition model is an advanced deep learning architecture that combines the advantages of Conformer networks in capturing local and global dependencies of speech signals with the characteristic of CTC (Connectionist Temporal Classification) in sequence labeling tasks that does not require explicit alignment. By fusing the local feature extraction capabilities of convolutional neural networks (CNNs) and the global dependency modeling capabilities of the Transformer's self-attention mechanism, Conformer networks can more effectively handle complex temporal information in speech signals. CTC, as the output layer, can directly decode text from the model's output sequence, simplifying the training process and improving recognition robustness. The end-to-end design means that the model can directly map from raw speech features to text output, reducing error accumulation in intermediate stages.
[0041] Specifically, the speech recognition model input is an 80-dimensional FBank feature, which is concatenated with the features of the preceding and following three frames to form an 80×7-dimensional input vector. FBank features (Mel filter bank features) are a commonly used acoustic feature in speech signal processing, effectively representing the spectral envelope information of speech, consistent with the auditory perception characteristics of the human ear. The 80-dimensional FBank feature provides rich frequency domain information. By concatenating the FBank features of the current frame with the features of the preceding and following three frames to form an 80×7-dimensional input vector, the model can be provided with broader contextual information, helping it better understand the dynamic changes in speech and the interactions between phonemes, thereby improving its ability to recognize phoneme boundaries and pronunciation details.
[0042] The encoder consists of 12 Conformer blocks. The encoder is the core of the speech recognition model, responsible for extracting high-level semantic information from the input features. Using 12 Conformer blocks to construct the encoder means the model has sufficient depth and complexity to learn the mapping relationship from low-level acoustic features to high-level linguistic features in the speech signal. Each Conformer block can perform multi-level abstraction and transformation of the input information, thereby capturing richer and more discriminative features.
[0043] Each Conformer block contains a feedforward module, a multi-head self-attention module, a convolutional module, and a second feedforward module. The feedforward module performs non-linear transformations on the features, increasing the model's expressive power. The multi-head self-attention module allows the model to focus on different parts of the input sequence in parallel across different "attention heads," thereby capturing dependencies at different time scales and semantic levels in the speech signal. This helps the model understand long-range contextual information, such as word associations. The convolutional module is specifically designed to capture local features and short-term dependencies in the speech signal, such as acoustic patterns within phonemes. Convolutional operations effectively extract local patterns and are robust to translation invariance in speech signals. The second feedforward module further performs non-linear mappings on the features processed by the attention mechanism and convolution, enhancing the expressive power of the features.
[0044] Based on this, the number of attention heads is 8, and the hidden unit dimension is 512. The 8 attention heads mean the model can simultaneously focus on the input sequence from 8 different angles or levels, thus capturing more comprehensive and detailed contextual information. The 512 hidden unit dimension determines the capacity of the model's internal representation; a larger dimension allows the model to store and process more complex feature information, thereby improving the model's learning ability and recognition accuracy.
[0045] The output layer is a Connectionist Temporal Classification (CTC) output layer. The CTC output layer is a loss function and decoding mechanism specifically designed for sequence-to-sequence tasks. It allows the model to directly predict the probability distribution of the output sequence without requiring precise pre-alignment of speech and text. CTC can handle the variable-length input and output sequence problems common in speech recognition and has good adaptability to changes in pronunciation speed, thus improving the robustness of recognition.
[0046] Furthermore, the vocabulary includes Chinese characters, numbers, letters, and automotive parts terminology. The vocabulary is a collection of all words that the speech recognition model can recognize. By incorporating Chinese characters, numbers, letters, and specialized automotive parts terminology into the vocabulary, the model is able to accurately recognize all instructions that operators might use in actual work, especially those specialized terms that may be uncommon or easily confused in general speech recognition models. This is crucial for the specialized application scenarios of the integrated platform for AI measurement and defect detection of automotive parts dimensions.
[0047] Through the above technical solution, the speech recognition unit adopts an end-to-end speech recognition model based on the Conformer-CTC architecture. This model combines the advantages of Conformer in capturing local and global dependencies in speech signals with the characteristic of CTC that it does not require explicit alignment in sequence annotation, significantly improving the accuracy and robustness of operator voice command recognition in complex industrial environments. Specifically, by using 80-dimensional FBank features concatenated with three frames before and after to form an 80×7-dimensional input vector, and incorporating multi-head self-attention and convolutional modules within the 12-layer Conformer block encoder, the model can more comprehensively and deeply understand the acoustic and temporal characteristics of the speech signal, effectively addressing environmental noise and pronunciation variations. Furthermore, a vocabulary specifically covering Chinese characters, numbers, letters, and automotive parts terminology ensures accurate recognition of professional commands, avoiding misoperations caused by terminology recognition errors. This enables the voice interaction module to more reliably convert operator voice commands into accurate control commands, thereby improving the automation level and operational efficiency of the entire platform, reducing the need for manual intervention, and ensuring the accurate execution of dimensional measurement and defect detection tasks.
[0048] In one alternative implementation, the Natural Language Understanding Unit employs a BERT-based joint model for intent classification and slot filling. This joint model is a deep learning model whose core idea is to simultaneously perform intent classification (understanding the user's overall dialogue purpose) and slot filling (extracting key information fragments from the user's statements). By training these two tasks together, the model can better capture the intrinsic correlation between intent and slots, thus exhibiting higher accuracy and robustness in understanding complex instructions. BERT (Bidirectional Encoder Representations from Transformers), as a pre-trained language model, can learn rich contextual semantic information, providing high-quality feature representations for intent classification and slot filling.
[0049] Specifically, the joint model uses the Chinese BERT-base as its backbone network. The Chinese BERT-base is a BERT model pre-trained on Chinese corpora, learning vocabulary, grammar, and semantic patterns from massive amounts of Chinese text data. Using it as the backbone network ensures the model's ability to understand Chinese instructions, especially when dealing with specialized terminology and expressions in the automotive parts field. It fully utilizes the Chinese language knowledge acquired during pre-training, thereby improving the model's performance in specific application scenarios. In terms of model structure, the output vector at the CLS position is input into a fully connected layer for intent classification. In the BERT model, a special classification label ([CLS]) is added to the beginning of each input sequence. After processing by the BERT encoder, the output vector corresponding to this [CLS] label is considered as the aggregate semantic representation of the entire input sequence. This [CLS] output vector is input into one or more fully connected layers, and through the Softmax activation function, it can be mapped to a predefined intent category space, thereby classifying the user's instruction intent. For example, determining whether the user wants to "start detection" or "query results". Simultaneously, the output vectors at each word position are input into a linear layer and a conditional random field for slot filling. In addition to the [CLS] token, BERT generates an output vector for each token in the input sequence, containing semantic information about that token in the current context. For slot filling, these token-level output vectors are first fed into a linear layer for feature transformation, and then further fed into a Conditional Random Field (CRF) layer. CRF is a sequence labeling model that considers dependencies between adjacent labels, thus exhibiting superior performance in slot boundary recognition and label assignment; for example, identifying "part type" as "engine block" and "measurement item" as "aperture".
[0050] The intent categories defined in this application include at least one of start detection, stop detection, query results, and set parameters. These intent categories represent the core operational instructions of the operator on the automotive parts inspection platform. For example, the "start detection" intent triggers the initiation of the inspection process; the "stop detection" intent is used to interrupt the current inspection; the "query results" intent is used to obtain the data or status of completed inspections; and the "set parameters" intent allows the operator to adjust the inspection threshold or mode. This clear classification of intents is the foundation for achieving precise control. Furthermore, the slot label set adopts the BIO annotation system, including at least one of part type, measurement item, value, and unit. The BIO (Begin, Inside, Outside) annotation system is a widely used annotation method in Named Entity Recognition (NER) tasks. BX represents the start term of entity X, IX represents the inside term of entity X, and O represents a non-entity term. Through this annotation system, the model can accurately identify key information fragments in the instructions, such as "part type" (e.g., "piston"), "measurement item" (e.g., "diameter"), "value" (e.g., "10.5"), and "unit" (e.g., "millimeters"), which are crucial for constructing specific control instructions.
[0051] Through the aforementioned technical solution, the Natural Language Understanding Unit (NLU) can leverage a BERT-based joint model, fully utilizing its powerful contextual understanding capabilities and Chinese language knowledge, to perform high-precision and robust intent classification and slot filling for operator voice commands. Specifically, by using the output vector at the CLS position for intent classification, the overall operational intent of the operator, such as start, stop, or query, can be accurately determined. Simultaneously, by combining the output vectors at each lexical position with a linear layer and a conditional random field for slot filling, key parameters such as part type, measurement item, value, and unit can be accurately extracted from the command. Even if the command contains technical terms or colloquial expressions, it can effectively avoid comprehension biases caused by semantic ambiguity or missing information. This joint modeling approach allows intent and slot information to mutually reinforce each other, improving the accuracy and consistency of overall understanding. Therefore, this application can significantly enhance the voice interaction module's ability to understand complex industrial commands, ensuring the accuracy and completeness of control command generation. This effectively solves the problems of inaccurate intent recognition and incomplete key parameter extraction in traditional methods when processing domain-specific voice commands, thereby improving the automation level and operational efficiency of the entire integrated platform for AI measurement and defect detection of automotive parts dimensions, and reducing the risk of manual intervention and potential misoperation.
[0052] In one alternative implementation, the dialogue management unit is configured as follows: Maintain the slot filling status table to record filled slots and slots to be filled; When all the necessary slots corresponding to the current intent output by the natural language understanding unit are filled, a control command is generated and sent to the decision control module. When there are unfilled required slots, the corresponding follow-up question template is selected from the preset follow-up question template library, the follow-up question text is generated, and the speech synthesis unit is triggered to broadcast it. The system receives the user's voice response to follow-up questions, processes it through the speech recognition unit and the natural language understanding unit, updates the slot filling status table, and repeats the above process until all necessary slots are filled or a cancellation command is received.
[0053] The slot filling status table is a core data structure within the dialogue management unit used to track and manage dialogue progress. This table is typically stored in key-value pairs, where the key represents the type of information to be collected (i.e., the slot, such as "part type," "measurement item," "value," etc.), and the value represents the current status of that slot (e.g., filled content, pending filling, optional, etc.). By maintaining this status table, the dialogue management unit can clearly understand the progress of the current task, identify which information has been successfully retrieved, and which information is still missing but necessary to complete the task. For example, when an operator issues a "start detection" command, the system initializes this status table based on the required slots associated with the preset "start detection" intent.
[0054] Once the slot filling status table shows that all necessary information corresponding to the current intent (i.e., all necessary slots) has been successfully extracted and filled from the operator's voice command, the dialogue management unit will determine that the task information is complete. At this point, the dialogue management unit will construct a structured control command based on this filled slot information, according to preset rules or templates. This control command is then sent to the decision control module to trigger the execution of subsequent hardware acquisition, data processing, or detection tasks. For example, if the "Part Type" slot of the "Start Detection" intent is filled with "Engine Block", the dialogue management unit will generate a control command containing "Start Detection" and "Part Type: Engine Block".
[0055] If the dialogue management unit detects one or more unfilled required slots for the current intent, it indicates that the current information is insufficient to perform the task. To proactively guide the operator to provide the missing information, the dialogue management unit intelligently selects the most suitable follow-up question template from a pre-built library, based on the type of the missing slot and the context. For example, if the "part type" slot is missing, the system might select the template "What type of part do you want to detect?" After selecting the template, the system generates the corresponding follow-up question text, converts it into a speech signal through the speech synthesis unit, and plays it to the operator, thus initiating a new round of interaction.
[0056] After a follow-up question is initiated, the dialogue management unit enters a state awaiting user response. When the operator responds to the follow-up question via voice, the voice response is first converted into text by the speech recognition unit, and then processed by the natural language understanding unit to extract new intent and slot information. The dialogue management unit then uses this new information to update the slot filling status table. This process is an iterative loop; the system continuously asks follow-up questions, receives responses, and updates the status until all necessary slots are successfully filled, or the operator explicitly issues an instruction to cancel the current task, at which point the dialogue flow terminates. This mechanism ensures that even through multiple rounds of interaction, the system can progressively collect all the information needed to complete the task.
[0057] Through the aforementioned technical solutions, the dialogue management unit can effectively manage the information flow during voice interaction. Even if the operator's initial instructions are incomplete, the system can proactively and purposefully guide the operator to provide all necessary detection parameters through follow-up questions. This multi-turn dialogue mechanism significantly improves the robustness and user experience of voice interaction, avoiding instruction execution failures or misoperations due to missing information. Specifically, by maintaining a slot filling status table, the system can clearly track the dialogue progress, ensuring that control instructions are generated only after all necessary information has been collected, thereby guaranteeing the accuracy and effectiveness of the instructions. Simultaneously, the pre-set follow-up question template library and iterative follow-up question-response mechanism enable the system to intelligently handle incomplete information, greatly improving the efficiency and intelligence level of voice interaction. This allows operators to interact with the platform in a more natural and flexible manner, thereby enhancing the overall usability and ease of operation of the integrated platform for AI measurement and defect detection of automotive parts dimensions.
[0058] In one alternative implementation, the dialogue management unit also has a built-in depth-based... Q A dialogue strategy optimization model for networks. This model takes the current slot filling state as input, the selection of follow-up question templates as actions, and the task success rate and the negative correlation between dialogue rounds as reward functions. It is trained in a simulated dialogue environment and is used to dynamically select the optimal follow-up question strategy in multi-turn dialogue scenarios.
[0059] Specifically, this dialogue strategy optimization model is a combination of deep learning and... Q The core of reinforcement learning algorithms lies in using neural networks to approximate... Q This function learns optimal dialogue strategies within a complex dialogue state space. The model learns which dialogue action (i.e., which follow-up question template) to take in a given dialogue state to maximize long-term rewards. In practice, this model typically includes a deep neural network whose input is a representation of the current dialogue state, and whose output is a function for each possible action. QDuring training, the model collects experiential data by interacting with a simulated dialogue environment and uses this data to update the weights of the neural network in order to gradually optimize its dialogue strategy.
[0060] The dialogue strategy optimization model uses the current slot filling state as input, enabling the model to fully perceive the current dialogue context. The slot filling state is the core representation of the dialogue management unit's understanding of user intent and information collection progress, including information on filled slots, slots yet to be filled, and user intent. This state information can be encoded into a vector or tensor, for example, using one-hot encoding to represent the slot filling state, and combined with the embedded representation of the specific values of filled slots to form a comprehensive input vector, thus providing the model with a decision-making basis.
[0061] Simultaneously, the selection of a follow-up question template is treated as an action. Each template in the preset follow-up question template library is assigned a unique identifier, serving as a discrete action in the model's action space. The model determines which follow-up question template to select by outputting the action with the highest Q-value, thus guiding the user to provide the required information.
[0062] To effectively guide model learning, this application uses a negative correlation between task success rate and dialogue rounds as the reward function. Task success rate serves as a positive reward, encouraging the model to efficiently complete information gathering tasks; the negative correlation between dialogue rounds (i.e., fewer dialogue rounds result in higher rewards) serves as a negative reward, prompting the model to learn strategies for maintaining concise dialogue. In a simulated dialogue environment, the model receives a positive reward when the dialogue is successfully completed (all necessary slots are filled and control commands are generated); for each dialogue round, the model receives a small negative reward to penalize excessively long dialogues. If the dialogue fails, a larger negative reward is given.
[0063] The model is trained in a simulated dialogue environment. Reinforcement learning models typically require large amounts of interactive data for training, while simulated dialogue environments can efficiently generate vast amounts of dialogue data, accelerating the model's learning process and ensuring its robustness before deployment. The simulated dialogue environment can mimic various user input behaviors and the responses of the dialogue management unit, updating the dialogue state and calculating rewards based on the model's actions and the simulated user's responses.
[0064] Through the aforementioned technical solution, the dialogue management unit no longer relies solely on preset fixed follow-up question templates, but can dynamically and intelligently select the optimal follow-up questioning strategy based on the real-time state of the current dialogue. This model learns in a simulated environment, using task success rate and dialogue rounds as rewards, enabling the system to balance the completeness of information acquisition with the efficiency of the dialogue. Specifically, when user input is incomplete, the model can predict which follow-up questioning method is most likely to quickly obtain the remaining key information based on the existing slot filling status, thereby avoiding redundant or inefficient questioning. This significantly improves the fluency and intelligence of multi-turn dialogues, reduces the number of interaction rounds between the operator and the system, improves the efficiency and accuracy of voice command input, and thus optimizes the overall user experience of the integrated platform for AI measurement and defect detection of automotive parts dimensions.
[0065] In one alternative implementation, the speech synthesis unit employs a cascaded architecture of the FastSpeech2 acoustic model and the HiFi-GAN vocoder; wherein the FastSpeech2 acoustic model takes a text sequence as input and generates a Mel spectrum through a duration predictor, a pitch predictor, and an energy predictor, while the HiFi-GAN vocoder takes the Mel spectrum as input and generates a speech waveform.
[0066] Specifically, the speech synthesis unit is a key component of the voice interaction module, its core function being to convert text information into audible speech signals. In automotive parts inspection platforms, it is responsible for clearly and accurately conveying complex inspection results, system status, or operator inquiries in speech form, making it a crucial link in achieving natural human-computer interaction. The FastSpeech2 acoustic model is a non-autoregressive text-to-mel-spectrogram generation model, whose main advantage lies in its ability to achieve fast and high-quality speech synthesis. This model takes a text sequence as input and explicitly models and controls the prosodic information of the speech through internal duration predictors, pitch predictors, and energy predictors. The duration predictor predicts the duration of each phoneme or word, the pitch predictor predicts the fundamental frequency (F0), and the energy predictor predicts the loudness. These predictors work together to enable the model to generate melodic and expressive melodic spectra, thus avoiding the slow inference speed and accumulated error problems inherent in traditional autoregressive models. The HiFi-GAN vocoder is a high-fidelity vocoder based on Generative Adversarial Networks (GANs). Its main function is to convert the Mel spectrum generated by the FastSpeech2 acoustic model into the original speech waveform. This vocoder is renowned for its superior sound quality and efficient inference speed, capable of generating speech waveforms highly similar to real human voices. Through adversarial training between the generator and multiple discriminators, it learns the complex mapping relationship from Mel spectrum to waveform, ensuring that the generated speech achieves a high level of timbre, clarity, and naturalness. The cascaded architecture refers to connecting the FastSpeech2 acoustic model and the HiFi-GAN vocoder in series to form a complete text-to-speech (TTS) system. In this architecture, the FastSpeech2 model first converts the input text sequence into a Mel spectrum, which is then used as input to the HiFi-GAN vocoder, which further generates the final speech waveform. This phased processing approach allows the system to fully leverage FastSpeech2's advantages in prosody control and Mel spectrum generation, as well as HiFi-GAN's advantages in waveform synthesis quality and efficiency, thereby optimizing the overall speech synthesis performance.
[0067] Through the aforementioned technical solution, the speech synthesis unit employs a cascaded architecture of the FastSpeech2 acoustic model and the HiFi-GAN vocoder, significantly improving the quality, naturalness, and efficiency of speech feedback. Specifically, the FastSpeech2 model processes the text sequence through duration predictors, pitch predictors, and energy predictors, precisely controlling the prosodic features of the speech to generate a richly expressive Mel spectrum, effectively avoiding the mechanical and unnatural phenomena that may occur in traditional speech synthesis. Subsequently, the HiFi-GAN vocoder uses this Mel spectrum as input to efficiently and faithfully generate the original speech waveform, ensuring the clarity and realism of the output speech. This cascaded architecture enables the system to quickly and accurately convert detection result data or follow-up text into high-quality, natural, and fluent speech signals, greatly improving the operator's auditory experience and enhancing the efficiency and comfort of human-computer interaction. Especially in industrial application scenarios such as the integrated AI measurement and defect detection platform for automotive parts, which requires real-time and accurate voice feedback, this solution can ensure that key information is conveyed to the operator in the clearest and most natural way, thereby improving the overall system's usability and ease of operation.
[0068] In one optional implementation, the voice interaction module further includes a communication interface unit; the communication interface unit is used to format the control commands generated by the dialogue management unit into JSON format data packets, send them to the decision control module via TCP / IP protocol, and receive JSON format detection result data returned by the decision control module, parse it, and send it to the speech synthesis unit for playback.
[0069] Specifically, the communication interface unit, as an independent communication entity, primarily manages the data flow between the voice interaction module and the decision control module. This unit is responsible for data formatting, encapsulation and decapsulation of transmission protocols, ensuring correct and efficient data transmission between different modules. In practical applications, the communication interface unit can be a standalone hardware module, such as a dedicated network communication chip or interface card, or a software module, such as a library or service implementing a specific communication protocol stack.
[0070] Once the dialogue management unit generates control commands, the communication interface unit formats these commands into JSON data packets. JSON (JavaScript Object Notation) is a lightweight data exchange format with a clear structure, easy for humans to read and write, and easy for machines to parse and generate. Encapsulating control commands in JSON format allows the command content to be presented in a structured, self-describing way, including fields such as "command type," "target part ID," and "detection parameters." This greatly simplifies the parsing and processing process for the receiver (i.e., the decision control module), improving the efficiency and accuracy of data exchange.
[0071] Subsequently, the communication interface unit sends the formatted JSON data packet to the decision control module via the TCP / IP protocol. TCP / IP is the foundational protocol suite of the Internet, providing reliable, connection-oriented data transmission services. Choosing TCP / IP as the underlying transmission mechanism ensures the integrity, order, and reliability of control command data packets during network transmission, preventing data loss or corruption, which is crucial for the accurate execution of control commands. The communication interface unit, acting as a client, establishes a TCP connection with the decision control module, which acts as a server, and transmits data through this connection.
[0072] After the decision control module completes the dimensional measurement and defect detection tasks, it generates corresponding inspection result data. To facilitate efficient and standardized data exchange with the voice interaction module, this result data is also encapsulated in JSON format and transmitted back via the communication interface unit. Upon receiving the JSON-formatted inspection result data from the decision control module, the communication interface unit parses it. This parsing process involves deserializing the JSON string into a structured data object, extracting key information to be broadcast, such as text content like "Inspection passed," "Dimensional deviation," and "Scratches found." Finally, this extracted text information is sent to the speech synthesis unit for conversion into speech for broadcast, thus providing voice feedback to the operator.
[0073] By introducing a communication interface unit and using JSON format data packets for data exchange with the TCP / IP protocol, this application effectively solves the standardization and reliability issues of data transmission between the voice interaction module and the decision control module. The JSON format makes the control commands and detection results data structure clear and easy to parse, greatly simplifying the design of data interfaces between different modules and reducing the difficulty of system integration. Simultaneously, the TCP / IP protocol ensures the integrity and order of data transmission, avoiding the risk of command loss or result corruption, thereby ensuring the accuracy of system control and the timeliness and reliability of detection result feedback. This standardized communication mechanism not only improves the modularity and maintainability of the entire platform but also provides a solid foundation for future system expansion and upgrades, enabling operators' voice commands to be executed accurately and detection results to be accurately fed back to the operator via voice.
[0074] In one alternative implementation, the multispectral imaging unit includes one or more of a visible light camera and an infrared thermal imager, used to acquire surface images of components under different spectral characteristics to highlight different types of texture defects.
[0075] Specifically, visible light cameras (such as color CCDs) acquire RGB images to detect common defects such as color differences, scratches, and dirt. Infrared thermal imagers (such as microbolometers) acquire thermal images, are sensitive to temperature distribution, and can be used to detect thermal anomalies caused by material inhomogeneity, internal cracks, or poor welding. Combining the two provides complementary information, improving the comprehensiveness of defect detection.
[0076] In one alternative implementation, the feature fusion module projects the 3D point cloud data onto the imaging plane of the multispectral 2D image data to achieve pixel-level data fusion, and inputs the fused data into a unified convolutional neural network to generate a joint feature vector in an end-to-end manner.
[0077] Specifically, the first step is to precisely calibrate the intrinsic and extrinsic parameters of the multispectral imaging unit and the high-precision 3D scanning unit to obtain their relative pose relationships and respective projection parameters. Then, using these calibration parameters, each 3D point cloud data point acquired by the high-precision 3D scanning unit is... Through perspective projection or orthographic projection algorithms, the corresponding pixel coordinates are accurately mapped onto the imaging plane of the multispectral two-dimensional image data. In this process, each projected pixel can be associated with its original 3D depth information, thus providing a foundation for subsequent fusion operations. Based on this, pixel-level data fusion is achieved, that is, at each pixel location in the image, the texture information of the 2D image is tightly combined with the corresponding 3D geometric information. This can be done by superimposing the projected 3D information (such as depth values, surface normals, or other local geometric features extracted from the point cloud) as an additional channel with the original channels of the multispectral 2D image (such as visible light R, G, B, infrared IR, etc.) to form a multi-channel fused data tensor. This fusion method ensures that each pixel contains not only its appearance information but also its precise spatial location and geometric morphology information. Subsequently, the fused data is input into a unified convolutional neural network. The unified convolutional neural network acts as an end-to-end learner, capable of simultaneously processing and understanding the fused multimodal data. The input layer of this CNN is designed to receive the aforementioned multi-channel fused data; for example, if the fused data contains 3 channels of the RGB image, 1 channel of the infrared image, and 1 channel of depth information, then the input to the CNN will be 5 channels. In this way, the network can jointly learn information from different modalities at an early stage, thereby capturing deep correlations and complementary features between them. This avoids the information loss that may occur when multimodal features are extracted independently and then simply concatenated in traditional methods. Finally, the joint feature vector is generated in an end-to-end manner. End-to-end learning means that the entire feature extraction process is a continuous, differentiable system, from the original fused data input to the final joint feature vector output, without the need for manual intervention in intermediate feature engineering. This unified convolutional neural network progressively extracts and abstracts the deep representation of the fused data through multiple layers of convolution, pooling, and activation operations. During training, the network is optimized according to the loss function of downstream tasks (such as dimensional measurement and defect detection), ensuring that the generated joint feature vector can maximally retain and express the two-dimensional texture features and three-dimensional geometric features of the parts, thus possessing high discriminative power for these tasks.
[0078] In one alternative implementation, the dimensional measurement task and the defect detection task in the integrated inspection module are executed in parallel on the same graphics processor through different task threads.
[0079] Specifically, on GPUs (such as the NVIDIA Tesla T4), CUDA streaming technology is used to distribute the computational tasks of dimensional measurement and defect detection into different streams for parallel execution. Dimensional measurement involves edge detection and geometric fitting, which can process multiple regions in parallel; defect detection involves forward propagation of convolutional networks, which can overlap with other computations. Through parallelization, the overall processing time is reduced, meeting the requirements for real-time online detection.
[0080] In one optional implementation, the integrated detection module further includes a result cross-checking submodule, which is used to compare the consistency of the intermediate geometric features of the dimensional measurement task with the three-dimensional structural defects identified by the defect detection task in terms of spatial location. If there is a logical contradiction between the two, the data re-sampling or re-detection process is triggered.
[0081] Specifically, intermediate geometric features refer to edge points, key points, or fitted baselines / surfaces extracted from dimensional measurements; 3D structural defects refer to defect areas (such as sets of indented points) output by defect detection. The result cross-checking submodule calculates the spatial overlap between the two. For example, if a defect detection report shows an indentation near an edge, and the indentation area covers the edge points used in the dimensional measurement, but the dimensional measurement result shows that the dimension is acceptable, there is a logical contradiction (the indentation should cause a change in dimension). In this case, a re-inspection or data re-acquisition is triggered to ensure the reliability of the detection.
[0082] In one optional implementation, the decision control module is also connected to a feedback optimization module. The feedback optimization module is used to perform correlation analysis on historical inspection data, corresponding component production batch information, and process parameters, and dynamically adjust the model parameters or inspection thresholds of the feature fusion module and the integrated inspection module.
[0083] The feedback optimization module collects long-term inspection data, including dimensional deviations, defect types, production batches, and process parameters (such as machine tool speed and feed rate). Through regression analysis or decision tree algorithms, it identifies the correlation between process parameters and quality indicators. For example, if a positive correlation is found between excessively large bore diameter and excessively high honing pressure, the module suggests adjusting the pressure to the process control system. Simultaneously, it fine-tunes the threshold or model parameters of the dimensional measurement network based on the new process conditions, achieving adaptive optimization.
[0084] In one optional implementation, the flexible transfer unit is equipped with a positioning and clamping device controlled by a programmable logic controller. The positioning and clamping device fine-tunes the posture of the parts based on the preliminary image analysis results acquired by the hardware acquisition module before acquisition, so that they are in the optimal inspection position.
[0085] Specifically, before entering the high-precision acquisition station, a low-resolution high-speed camera is used to photograph the component for contour matching or feature point detection, determining the deviation between the current posture and the standard posture (such as rotation angle or translation offset). The programmable logic controller (PLC) controls the gripper or rotary table to make fine adjustments based on the deviation, ensuring that the component is in the optimal detection position and improving the accuracy of subsequent acquisition and measurement.
[0086] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. An integrated platform for AI measurement and defect detection of automotive parts dimensions, characterized in that, include: The voice interaction module is used to receive and recognize the operator's voice commands, convert the voice commands into control commands, and feed back the detection results to the operator through voice synthesis. The hardware acquisition module, connected to the voice interaction module, responds to the acquisition start command in the control command and includes a multispectral imaging unit, a high-precision three-dimensional scanning unit and a flexible transmission unit. The multispectral imaging unit and the high-precision three-dimensional scanning unit are arranged sequentially or collaboratively along the transmission direction of the flexible transmission unit, and are used to acquire multispectral two-dimensional image data and three-dimensional point cloud data of automotive parts synchronously or in time-division. The feature fusion module, connected to the hardware acquisition module, is used to align the multispectral two-dimensional image data and the three-dimensional point cloud data of the same automotive part in terms of spatial coordinates, and to extract a joint feature vector that fuses two-dimensional texture features and three-dimensional geometric features based on a deep learning network. An integrated detection module, connected to the feature fusion module, has a built-in shared joint feature vector processing channel for parallel or pipelined execution. The dimensional measurement task, based on the three-dimensional geometric feature information in the joint feature vector, calculates the key dimensional parameters of the parts through sub-pixel level edge detection and geometric fitting algorithms; The defect detection task, based on the two-dimensional texture features and three-dimensional geometric features in the joint feature vector, identifies texture defects and three-dimensional structural defects on the surface of parts through a pre-trained defect classification and segmentation model. The decision control module, connected to the integrated detection module and the voice interaction module, is used to generate detection result data based on the combined results of the size measurement task and the defect detection task, and send the detection result data to the voice interaction module for voice feedback. At the same time, it generates sorting control instructions for the parts on the flexible conveying unit based on the detection result data.
2. The platform according to claim 1, characterized in that, The voice interaction module includes: The voice acquisition unit is used to acquire the raw voice signals emitted by the operator. The preprocessing unit, connected to the speech acquisition unit, is used to perform echo cancellation, noise suppression, beamforming and speech activity detection on the original speech signal, and output an enhanced speech segment. A speech recognition unit, connected to the preprocessing unit, is used to convert the enhanced speech segment into a text instruction; The natural language understanding unit, connected to the speech recognition unit, is used to perform intent classification and slot filling on the text instructions and extract at least one of the detection object identifier, detection item, and detection parameters. The dialogue management unit, connected to the natural language understanding unit, is used to maintain the dialogue state and, based on the matching status of the currently filled slots and the required slots, decide whether to generate control commands or initiate follow-up questions. The speech synthesis unit, connected to the dialogue management unit and the decision control module, is used to convert the detection result data or follow-up text into speech signals and play them.
3. The platform according to claim 2, characterized in that, The preprocessing unit is configured as follows: Echo cancellation is performed using a normalized minimum mean square adaptive filter; Noise suppression is achieved using spectral subtraction based on logarithmic minimum mean square error; A time difference of arrival estimation algorithm based on generalized cross-correlation-phase transform is used to locate the sound source, and a multi-channel signal is weighted and summed based on a delay summation beamformer, with the beam pointing in the direction of the sound source.
4. The platform according to claim 2, characterized in that, The speech recognition unit adopts an end-to-end speech recognition model based on the Conformer-CTC architecture; The input to the speech recognition model is 80-dimensional FBank features. The input vector is formed by concatenating three frames before and after the input. The encoder is a 12-layer Conformer block. Each Conformer block contains a feedforward module, a multi-head self-attention module, a convolution module, and a second feedforward module. The number of attention heads is 8, the hidden unit dimension is 512, and the output layer is a connection-series classification output layer. The vocabulary covers Chinese characters, numbers, letters, and automotive parts professional terms.
5. The platform according to claim 2, characterized in that, The natural language understanding unit uses a joint model based on BERT for intent classification and slot filling. The joint model uses Chinese BERT-base as the backbone network, inputs the output vector at the CLS position into a fully connected layer for intent classification, and inputs the output vector at each word position into a linear layer and a conditional random field for slot filling. The intent categories include at least one of start detection, stop detection, query result, and setting parameters. The slot label set adopts the BIO labeling system, including at least one of part type, measurement item, value, and unit.
6. The platform according to claim 2, characterized in that, The dialogue management unit is configured as follows: Maintain the slot filling status table to record filled slots and slots to be filled; When all the required slots corresponding to the current intent output by the natural language understanding unit are filled, the control command is generated and sent to the decision control module. When there are unfilled required slots, the corresponding follow-up question template is selected from the preset follow-up question template library, the follow-up question text is generated, and the speech synthesis unit is triggered to broadcast it. The system receives the user's voice response to follow-up questions, processes it through the speech recognition unit and the natural language understanding unit, updates the slot filling status table, and repeats the above process until all necessary slots are filled or a cancellation command is received.
7. The platform according to claim 6, characterized in that, The dialogue management unit also has a built-in depth-based function. Q A dialogue strategy optimization model for the network, wherein the current slot filling state is taken as the input state, the follow-up question template selection is taken as the action, and the task success rate and the negative correlation between the number of dialogue rounds are taken as the reward function, is trained in a simulated dialogue environment, and is used to dynamically select the optimal follow-up question strategy in multi-turn dialogue scenarios.
8. The platform according to claim 2, characterized in that, The speech synthesis unit adopts a cascaded architecture of the FastSpeech2 acoustic model and the HiFi-GAN vocoder. The FastSpeech2 acoustic model takes a text sequence as input and generates a Mel spectrum through a duration predictor, a pitch predictor, and an energy predictor. The HiFi-GAN vocoder takes the Mel spectrum as input and generates a speech waveform.
9. The platform according to claim 2, characterized in that, The voice interaction module further includes a communication interface unit; the communication interface unit is used to format the control commands generated by the dialogue management unit into JSON format data packets, send them to the decision control module via TCP / IP protocol, and receive the JSON format detection result data returned by the decision control module, parse it, and send it to the speech synthesis unit for broadcast.