Vehicle repair store member extension method and device based on AI, and electronic equipment

By collecting multimodal data in auto repair stores and performing feature fusion, and using AI models to generate member customer expansion strategies, the problem of low manual statistics efficiency is solved, and accurate member customer expansion and customer loyalty is achieved.

CN120338880AInactive Publication Date: 2025-07-18DECHE CHUANGRONG (BEIJING) TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510483384.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Auto repair stores rely on manual statistical methods in the process of member customer acquisition, which makes data processing time-consuming and labor-intensive, difficult to quickly capture changes in customer behavior, and difficult to formulate accurate member customer acquisition strategies.

Method used

Using AI-based methods, multimodal data (text, image, time series data), feature fusion, member customer expansion strategies are generated, and accurate member customer expansion strategies are automatically generated using AI models.

Benefits of technology

It has achieved a multi-dimensional understanding of customers and store conditions, improved processing speed and customer expansion efficiency, customized customer expansion solutions, and improved member conversion rates and customer loyalty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338880A_ABST
    Figure CN120338880A_ABST
Patent Text Reader

Abstract

The invention provides an AI-based vehicle repair store member extension method and apparatus, and an electronic device, and relates to the field of data processing. In the method, multi-modal data for a target automobile repair store and a target customer group is acquired, the multi-modal data comprises text data, image data and time sequence data, the text data comprises customer evaluation and social media information, the image data comprises automobile condition photos and store environments, and the time sequence data comprises maintenance records and consumption tracks; performing feature fusion on the multi-modal data to obtain a multi-modal feature group; and inputting the multi-modal feature group into an AI model to generate a member customer extension strategy for the target customer group so as to perform member customer extension on the target automobile repair store. By implementing the technical scheme provided by the invention, an accurate member extension strategy can be conveniently generated and formulated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to an AI-based method, device, and electronic equipment for attracting members in auto repair shops. Background Art

[0002] With the continuous growth of the automobile ownership and the continuous upgrading of consumers' demand for after-sales services in the automotive market, auto repair shops are facing unprecedented opportunities and challenges. In a highly competitive market environment, consumers are paying more and more attention to brand experience and service quality, which prompts auto repair shops to transform from single repair services to comprehensive after-sales services in the automotive market. Against this background, attracting members has become an important way for auto repair shops to enhance customer loyalty and increase additional service revenue.

[0003] Currently, auto repair shops mainly rely on manual statistical methods in the process of attracting members. This method manually sorts out the repair records of the shop, customer consumption data, and member registration information, and uses tools such as Excel for simple data aggregation and calculation. However, due to the scattered data sources and low update frequency, manual processing is not only time-consuming and laborious, but also difficult to quickly capture the subtle changes in customer behavior when the data volume is large or the data dimension is numerous, resulting in the difficulty of formulating accurate member attraction strategies.

[0004] Therefore, there is an urgent need for an AI-based method, device, and electronic equipment for attracting members in auto repair shops. Summary of the Invention

[0005] This application provides an AI-based method, device, and electronic equipment for attracting members in auto repair shops, which is convenient for generating and formulating accurate member attraction strategies.

[0006] In the first aspect of this application, an AI-based method for attracting members in auto repair shops is provided. The method includes: obtaining multimodal data for a target auto repair shop and a target customer group, where the multimodal data includes text data, image data, and time series data. The text data includes customer evaluations and social media information, the image data includes vehicle condition photos and shop environment, and the time series data includes repair records and consumption trajectories; performing feature fusion on the multimodal data to obtain a multimodal feature group; inputting the multimodal feature group into an AI model to generate a member attraction strategy for the target customer group, so as to attract members for the target auto repair shop.

[0007] By adopting the above technical solutions, by collecting text, image and time-series data, a comprehensive understanding of customers and store conditions has been achieved from multiple dimensions. Feature fusion of multi-modal data can integrate the advantages of each data source, extract richer and more accurate customer and store features, and reduce the bias brought by a single data source. Inputting the fused features into the AI model to automatically generate membership customer acquisition strategies for the target customer group can quickly respond to market changes and improve the scientificity and accuracy of decision-making. The entire process, from data collection, feature fusion to strategy generation, is automated, greatly reducing the workload of manual statistics, and improving the processing speed and customer acquisition efficiency. Based on multi-modal data and AI analysis, a customer acquisition plan can be customized for each target auto repair store, helping the store to more effectively acquire high-quality customers and improve the membership conversion rate and customer loyalty. Therefore, it is convenient to generate accurate membership customer acquisition strategies.

[0008] Optionally, the obtaining of multi-modal data for the target auto repair store and the target customer group specifically includes: receiving the original data sent by the store terminal device for the target auto repair store and the target customer group; performing word segmentation, cleaning, and standardization processing on the data related to the text part in the original data to obtain the text data; performing normalization and data augmentation on the data related to the image part in the original data to obtain the image data; performing smoothing, normalization, and time window segmentation on the data related to the time series part in the original data to obtain the time series data; generating the multi-modal data according to the text data, the image data, and the time series data.

[0009] By adopting the above technical solutions, collecting text, image and time-series data simultaneously ensures a multi-dimensional and comprehensive understanding of the target auto repair store and the customer group, avoiding the problem of one-sided information brought by a single data source. Adopting specialized preprocessing methods for different data types (text word segmentation and cleaning, image normalization and enhancement, time series data smoothing and time window segmentation) can improve data quality, reduce noise, and ensure the accuracy and stability of subsequent processing and model input. By performing standardization processing on the original data, various data are easier to fuse after unified processing, laying a good foundation for constructing an efficient and intelligent multi-modal model in the future. The entire process from data collection to preprocessing and then to multi-modal data generation realizes automated processing, reduces manual intervention and statistical workload, and improves the overall processing efficiency and response speed. After the multi-modal data is generated, it can combine the complementary advantages of various types of information to provide richer and more accurate input for the subsequent AI model, helping to generate more accurate membership customer acquisition strategies.

[0010] Optionally, when fusing the features of the multi-modal data to obtain a multi-modal feature group, the following formula is specifically used for calculation: ; Among them, h fused is a multi-modal feature group, LayerNorm is used to normalize the fused feature vector to enhance the model stability and convergence speed, M is the number of modalities, m is the current modality, g (m) is the gating coefficient, A (m) is the attention output, is the element-wise multiplication, which is used to multiply g (m) with A (m) to achieve dynamic weighting for each dimension.

[0011] By adopting the above technical solutions, by fusing the features of different modalities, the complementary information in text, images, and time-series data can be integrated together, thereby obtaining a richer and more comprehensive customer feature representation. The element-wise multiplication of the gating coefficient and the attention output realizes dynamic weighting for each dimension, effectively highlighting key features and reducing the influence of noise. Using LayerNorm to normalize the fused feature vector can enhance the stability of the model and accelerate the convergence speed during training. By managing the number of modalities and fusing them one by one, it can not only adapt to different types of data but also flexibly adjust the weights of each modality according to specific tasks, improving the overall performance of the model.

[0012] Optionally, the g (m) is specifically calculated using the following formula: ; Among them, represents the Sigmoid activation function, which is used to map the output value between 0 and 1 as the gating factor for this modality, is the gating weight matrix, h (m) is the input feature vector representing the m-th modality, is the bias term vector.

[0013] By adopting the above technical solutions, by using the Sigmoid activation function to map the output value between 0 and 1, a reasonable gating factor can be generated for each modality, thereby adaptively adjusting the influence of each modality in feature fusion. Using the gating weight matrix and the bias term to perform a linear transformation on the input features enables the model to automatically learn and emphasize more valuable information during training and filter out noise. The dynamic gating mechanism ensures that during multi-modal data fusion, the contributions of each modality can be dynamically adjusted according to the data quality, thereby generating more stable and distinguishable fused features. Through normalization and dynamic weighting processing, the model can converge faster and maintain high stability and accuracy during training and inference.

[0014] Optionally, the A (m)The specific calculation is carried out using the following formula: ; Among them, Q (m) is the query of mode m after linear transformation, k is all modes, is used to represent calculating the attention weights of the query of each mode m and the keys of all modes k to obtain a probability distribution, indicating how mode m affects the information of other modes.

[0015] By adopting the above technical solution, by matching the query of the mode with the keys of all modes k, the information association between each mode can be fully captured, and the effective complementarity of cross-modal data can be realized. The probability distribution generated by the attention mechanism can automatically assign weights to the interaction between each mode, enabling important information to receive higher attention, thereby improving the quality of the fused features. This formula enables each mode to not only rely on its own information but also dynamically absorb useful information from other modes, contributing to constructing a more comprehensive and robust multi-modal feature representation. By calculating cross-modal attention, the model can better understand the relationship between different modes, thus showing stronger performance in the subsequent generation of membership customer acquisition strategies.

[0016] Optionally, when inputting the multi-modal feature group into the AI model to generate a membership customer acquisition strategy for the target customer group, the specific calculation is carried out using the following formula: ; Among them, is the membership customer acquisition strategy, used to represent the probability of converting customers in the target customer group into members, is the logit value generated by the neural network part, that is, the conversion tendency describing the customer's own characteristics, is the weighted term of the distance part, used to represent the similarity between the customer and the customer group with high conversion tendency, and λ is used to adjust the contribution of the distance score in the overall logit.

[0017] By adopting the above technical solution, the conversion tendency of the customer himself (the logit value generated by the neural network) is combined with the similarity between the customer and the customer group with high conversion tendency (the distance weighted term), so as to more comprehensively capture the conversion potential of the customer. By introducing the parameter λ, the influence of the distance score in the overall logit can be flexibly adjusted, enabling the model to balance the importance of individual characteristics and group similarity in different scenarios. Considering the internal characteristics and similarity information comprehensively makes the generated membership customer acquisition strategy more robust and accurate, contributing to accurately locking high-potential customers.

[0018] Optionally, the method further includes: extracting a target probability value from the membership customer acquisition strategy; comparing the target probability value with a preset probability threshold; if it is determined that the preset condition indicates that the target probability value is greater than or equal to the target probability threshold, then performing membership customer acquisition on the customers in the target customer group who meet the preset condition.

[0019] By adopting the above technical solution, by extracting the target probability value and comparing it with the preset threshold, customer acquisition is only performed on customers who meet the conditions, thereby ensuring that marketing resources are concentrated on high-potential customers. Judging based on the probability value output by the model reduces subjective factors and improves the scientificity and objectivity of decision-making. The automated screening mechanism avoids blind manual delivery, can quickly identify customers who meet the preset conditions, and realizes precision marketing. Only performing customer acquisition on customers with a conversion probability higher than the threshold helps to reduce marketing costs and improve the input-output ratio.

[0020] In a second aspect of the present application, there is provided an AI-based membership customer acquisition device for an auto repair shop. The membership customer acquisition device for the auto repair shop includes an acquisition module and a processing module. Among them, the acquisition module is used to acquire multimodal data for a target auto repair shop and a target customer group. The multimodal data includes text data, image data, and time series data. The text data includes customer evaluations and social media information. The image data includes vehicle condition photos and store environments. The time series data includes repair records and consumption trajectories. The processing module is used to perform feature fusion on the multimodal data to obtain a multimodal feature group. The processing module is further used to input the multimodal feature group into an AI model to generate a membership customer acquisition strategy for the target customer group to perform membership customer acquisition on the target auto repair shop.

[0021] In a third aspect of the present application, there is provided an electronic device. The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described above.

[0022] In a fourth aspect of the present application, there is provided a computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed, execute the method described above.

[0023] In summary, one or more technical solutions provided in the present application have at least the following technical effects or advantages: By collecting text, images, and time-series data, a comprehensive understanding of customers and store conditions has been achieved from multiple dimensions. Feature fusion of multi-modal data can integrate the advantages of each data source, extract richer and more accurate customer and store features, and reduce the bias brought by a single data source. Inputting the fused features into an AI model to automatically generate membership acquisition strategies for the target customer group can quickly respond to market changes, improve the scientificity and accuracy of decision-making. The entire process, from data collection, feature fusion to strategy generation, is automated, greatly reducing the workload of manual statistics and improving the processing speed and customer acquisition efficiency. Based on multi-modal data and AI analysis, a customer acquisition plan can be customized for each target auto repair store, helping the store more effectively acquire high-quality customers and improve member conversion rates and customer loyalty. Therefore, it is convenient to generate precise membership acquisition strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 FIG. is a schematic flowchart of a method for acquiring members of an auto repair store based on AI provided by an embodiment of the present application; Figure 2 FIG. is another schematic flowchart of a method for acquiring members of an auto repair store based on AI provided by an embodiment of the present application; Figure 3 FIG. is a schematic block diagram of a device for acquiring members of an auto repair store based on AI provided by an embodiment of the present application; Figure 4 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0025] Description of reference numerals: 31, acquisition module; 32, processing module; 41, processor; 42, communication bus; 43, user interface; 44, network interface; 45, memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0027] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for instance" is intended to present related concepts in a specific manner.

[0028] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0029] With the continuous increase in the number of cars and the continuous upgrading of consumers' demand for aftermarket services, auto repair shops are facing unprecedented opportunities, but also fierce competition and severe challenges. In the fierce market competition, consumers' requirements for brand experience and service quality are increasing, prompting auto repair shops to transform from single maintenance services to all-round and comprehensive aftermarket services. In this context, member acquisition has become an important means to improve customer loyalty and increase additional service revenue.

[0030] At present, most auto repair shops still rely mainly on manual statistical methods in the process of member acquisition. This method manually organizes store maintenance records, customer consumption data and member registration information, and uses tools such as Excel to perform simple data aggregation and calculations. However, due to the scattered data sources and low update frequency, manual processing is not only time-consuming and laborious, but also difficult to quickly capture subtle changes in customer behavior when faced with large amounts of data or multi-dimensional data, making it difficult to formulate accurate member acquisition strategies.

[0031] In order to solve the above technical problems, this application provides an AI-based auto repair shop membership expansion method. Figure 1 , Figure 1 A flowchart of an AI-based auto repair shop member acquisition method provided in an embodiment of the present application. The method is applied to a server and includes steps S110 to S130, which are as follows: S110. Acquire multimodal data for target auto repair shops and target customer groups. The multimodal data includes text data, image data, and time series data. The text data includes customer reviews and social media information. The image data includes photos of vehicle conditions and store environments. The time series data includes maintenance records and consumption trajectories.

[0032] Specifically, a server refers to a computer system or a cloud computing platform used for collecting, storing, processing, and analyzing data. It plays a core role in data processing and decision-making support during the membership acquisition process of an auto repair shop. Multimodal data refers to the fusion of information from multiple different types of data sources. In this scenario, text data includes customer feedback, social media interactions, etc. recorded in text form; image data includes the vehicle status, store environment, etc. captured through photos or videos; time-series data includes customer repair records, consumption behavior trajectories, etc. that change over time. These data types complement each other and can depict customers' behaviors and needs from different dimensions, improving the accuracy of customer acquisition strategies. Combining text, image, and time-series data can provide a more comprehensive analysis of customers and store operations.

[0033] Among them, the disadvantages of single data are as follows: For example, relying solely on text data may not be able to determine customers' actual vehicle repair needs. Relying solely on image data, it is impossible to understand customers' subjective experiences of store services. Relying solely on time-series data, it is difficult to predict customers' brand loyalty and word-of-mouth dissemination ability. The value of fused data is as follows: Analyze customers' subjective satisfaction through text data, analyze the actual vehicle repair needs through image data, predict customers' next repair time through time-series data. Combining these data can generate more accurate membership acquisition strategies, such as: sending brake maintenance discount information to customers who frequently search for "brake repair" recently; pushing exclusive follow-up offers to old customers who have not visited the store in the past 3 months.

[0034] In a possible implementation manner, obtaining multimodal data for a target auto repair shop and a target customer group specifically includes: receiving the original data sent by the store terminal device for the target auto repair shop and the target customer group; performing word segmentation, cleaning, and standardization processing on the data related to the text part in the original data to obtain text data; performing normalization and data augmentation on the data related to the image part in the original data to obtain image data; performing smoothing, normalization, and time window segmentation on the data related to the time-series part in the original data to obtain time-series data; generating multimodal data based on the text data, image data, and time-series data.

[0035] Specifically, the server receives the original data from the store terminal devices (such as customer management systems, surveillance cameras, repair equipment, POS machines, etc.). This data may be structured (such as numerical information) or unstructured (such as text, pictures). For example, the store management system uploads customer repair records and consumption data. The surveillance camera captures vehicle condition photos and store environment images. The POS machine records customer payment information and generates a consumption trajectory. The customer evaluation system collects customers' evaluations on the store App or social media.

[0036] Among them, the text data mainly includes customer evaluations, social media information, etc. These data need to be tokenized, cleaned, and normalized for analyzing customer sentiment and needs. Tokenization is to split customer reviews into individual words or phrases for analysis. Cleaning is used to remove meaningless characters (such as emojis, special symbols), deduplicate, and remove stop words (such as "de", "shi"). Normalization is used to normalize synonyms, for example, "change engine oil" and "replace lubricating oil" are classified into the same category. The image data mainly includes vehicle condition photos and store environment photos, which need to be normalized and data augmented to improve the accuracy of AI recognition. Normalization is used to adjust image brightness, contrast, and unify the resolution to make the data format consistent. Data augmentation includes rotation, cropping, flipping, etc., to expand the dataset and improve the generalization ability of the model. The time series data refers to repair records and consumption trajectories, which need to be smoothed, normalized, and time window sliced to mine customer consumption trends. Smoothing is used to remove outliers, such as the deviation caused by a one-time large consumption. Normalization can standardize data such as consumption amount and repair frequency to avoid the influence of different numerical scales on analysis. Time window slicing divides the data by day, week, or month to analyze the long-term behavior patterns of customers. These data can be used as the input of the AI model to support precise member acquisition strategies.

[0037] S120. Perform feature fusion on the multi-modal data to obtain a multi-modal feature group.

[0038] Specifically, feature fusion refers to converting, aligning, and merging data of different modalities so that they can be uniformly processed by machine learning or deep learning models. In the embodiments of this application, the server receives text data, image data, and time series data, then extracts, transforms, and fuses these data into a multi-modal feature group, which is finally used for customer analysis and member acquisition strategy formulation.

[0039] Among them, data of different modalities need to be first feature extracted and converted into numerical vectors for subsequent fusion. For example: text data (customer evaluations, social media information) uses methods such as BERT, TF-IDF, and word vectors to extract semantic features. Image data (vehicle condition photos, store environment) uses convolutional neural network (CNN) to extract features such as vehicle damage areas and store cleanliness. Time series data (repair records, consumption trajectories) uses long short-term memory network (LSTM) or statistical methods to extract features such as consumption cycles and repair frequencies. In addition, the feature dimensions and scales of different modalities are different, so alignment is required. Normalize the numerical ranges of different features to make them on the same scale. Align the time steps of different modalities. For example, divide the time series data by month or quarter to synchronize it with text and image data. The fusion methods include concatenation, weighted average, gating mechanism, and attention mechanism.

[0040] In a possible implementation, feature fusion is performed on multi-modal data to obtain a multi-modal feature group, and the following formula is specifically used for calculation: ; where h fused is the multi-modal feature group, LayerNorm is used to normalize the fused feature vectors to enhance the model stability and convergence speed, M is the number of modalities, m is the current modality, g (m) is the gating coefficient, A (m) is the attention output, is the element-wise multiplication, which is used to multiply g (m) by A (m) to achieve dynamic weighting for each dimension.

[0041] Specifically, this formula describes how to fuse the features of multiple modalities (such as text, images, time-series data, etc.) to generate a comprehensive feature for subsequent analysis and decision-making. The idea is to control the contribution of each modality through the gating coefficient g (m) to determine which modalities are more important. The mutual influence between modalities is further measured through the attention weight A (m) . The weighted multi-modal feature group h fused is calculated. Each modality has its own feature vector. For example: text features extracted from the text modality (such as semantic information of customer reviews). Visual features such as the degree of vehicle damage and store environment extracted from the image modality. Time-series features such as the customer's historical repair records and consumption trajectories extracted from the time-series modality. The gating coefficient is used to measure the importance of different modalities, similar to a filtering mechanism: for example, if the customer's consumption records can predict membership conversion better than social media comments, then the time-series modality will contribute more than the text modality. The attention weight is used to measure how the current modality affects other modalities. For example, if a customer's review mentions "affordable price" and his consumption records also show that he often chooses discounted services, the attention mechanism will increase the weight of the text modality.

[0042] For example, assume a customer's data is as follows: Text data (customer review): "The service in this store is very good and the price is fair." h text =[0.8, 0.6, 0.2]; Image data (car condition photos): The car has minor scratches, h image =[0.3, 0.7, 0.6]; Time-series data (consumption records): Consumption amount in the past six months: [300 yuan, 500 yuan, 0 yuan, 700 yuan, 600 yuan, 0 yuan], h time =[0.5, 0.4, 0.8]. The server calculates the gating coefficient and attention weight, g text =0.7, g image =0.5, g time= 0.8, A text = 0.6, A image = 0.4, A time = 0.9. Calculate multimodal fusion: h fused =(0.7 × 0.6 × [0.8, 0.6, 0.2])+(0.5 × 0.4 × [0.3, 0.7, 0.6])+(0.8 × 0.9 × [0.5, 0.4, 0.8]), to obtain h fused =[0.62, 0.55, 0.68]. This fused feature vector can be used to predict the membership conversion probability of customers.

[0043] In a possible implementation, g (m) Specifically, it is calculated using the following formula: ; where represents the Sigmoid activation function, which is used to map the output value between 0 and 1 as the gating factor for this modality, is the gating weight matrix, h (m) is the input feature vector representing the m-th modality, is the bias term vector.

[0044] Specifically, the Sigmoid activation function is a non-linear transformation function, and the output value of this function is between 0 and 1. Therefore, it is particularly suitable for the gating mechanism, that is, mapping the contribution of a specific modality into a weight factor to control the influence of this modality in the fused feature. Since its value is always between (0, 1), it is suitable as a weight, avoiding the problem of gradient explosion or gradient disappearance caused by extreme values. The derivative of the function changes continuously, ensuring that the model can perform effective gradient optimization through backpropagation, thereby enhancing the learning ability. During the model training process, the output of Sigmoid can prevent the modality weights from overly biasing towards a certain modality, helping to optimize the problem of unbalanced data distribution.

[0045] In the process of multimodal data fusion, the main goal of the gating mechanism is to adaptively adjust the contribution weights of each modality, so that the features of important modalities are strengthened while the influence of noise modalities is weakened. The gating factor of modality m, ranging from 0 to 1, is used to measure the contribution degree of this modality to the fused feature. The gating weight matrix learns the weight distribution of different modalities through training. The bias term helps to adjust the feature distribution and improve the learning ability of the model.

[0046] For example, assume that it is necessary to analyze the customer behavior of an auto repair shop to predict which customers are most likely to become loyal members. The available data includes: Text data: Customer reviews on social media, such as "The service attitude of this shop is very good, and I will come again next time" or "The repair speed is too slow, not recommended." Image data: Vehicle repair photos uploaded by customers, reflecting the degree of vehicle damage and repair quality. Time-series data: Customers' past consumption records, including the frequency of entering the store, repair categories, consumption amounts, etc. In traditional methods, these data may be simply concatenated or averaged and weighted, resulting in information redundancy or noise interference. For example, if a customer's social media review is relatively neutral, but their consumption records show that they often visit the store, then relying solely on text data may not accurately determine their loyalty. Through the Sigmoid gating mechanism, the importance of different modalities can be adaptively adjusted. For example: If the text data contains extremely positive or negative reviews, the gating factor will be close to 1, emphasizing the role of the text modality in feature fusion. If the quality of the image data is poor (blurry or lacking key information), the gating factor will be close to 0, weakening its contribution. If the customer's consumption data indicates that they are a high-frequency consumer, the gating factor will be higher, ensuring that the time-series data contributes more to the final prediction. In this way, the model can automatically adjust the weights of different modalities of data to ensure that the final fused features can accurately reflect the customer behavior characteristics.

[0047] Therefore, the Sigmoid gating mechanism can automatically adjust the weights of each modality according to the data quality, ensuring that high-quality data obtains higher weights while the impact of low-quality data is weakened. By normalizing the weights of different modalities, the gating mechanism can reduce data bias during feature fusion, enabling the model to converge faster during the training process and avoiding problems such as gradient explosion or gradient disappearance. During the inference stage, the gating mechanism can effectively reduce the interference of noisy data on the final prediction, thereby improving the stability and generalization ability of the model. Even if some modality of data is missing (for example, a customer does not upload a repair photo), the model can still make reasonable predictions based on the data of other modalities.

[0048] In a possible implementation manner, ; where Q (m) is the query of modality m after linear transformation, k is all modalities, is used to represent calculating the attention weights between the query of each modality m and the keys of all modalities k to obtain a probability distribution, indicating how modality m affects the information of other modalities.

[0049] Specifically, traditional single-modal analysis often makes decisions based on only one type of data. For example, text analysis can use information such as customer evaluations and social media interaction records to judge customers' satisfaction with the store and service needs; image analysis can extract key information from data such as vehicle condition photos and repair record pictures to evaluate customers' vehicle repair needs; time-series data analysis can use data such as customers' repair history and consumption trajectories to predict customers' future consumption behavior. However, single-modal data often has limitations. For example, relying solely on text analysis may not accurately identify customers' current repair needs, and relying solely on time-series data may not evaluate customers' perception of the store service quality. Therefore, cross-modal data fusion has become a more effective solution, which can jointly model different types of data to improve the accuracy of customer behavior prediction.

[0050] In the embodiments of the present application, the idea of the cross-modal attention mechanism is to calculate the attention weights between different modalities, so that the model can automatically learn the mutual influence relationship between each modality and perform dynamic weighting according to its importance. This process mainly includes the following steps: for each modality m, first generate a query vector Q through a linear transformation. (m) a key vector K (m) and a value vector V (m) . Among them, the query vector represents the information requirement that the current modality hopes to interact with other modalities, the key vector represents the feature representation of all modalities, and the value vector represents the corresponding modality content information. By calculating the similarity between the query and the keys of all modalities, the cross-modal attention weights are obtained. Among them, d a is the dimension of the key vector, which is used for scaling normalization to avoid too large calculation results. This attention weight indicates how modality m is affected by modality k. The larger the weight, the stronger the correlation between the two modalities. Use the calculated attention weights to perform weighted summation on the value vectors. In this way, each modality not only depends on its own information, but can dynamically absorb data from other modalities, thereby forming a more comprehensive fused feature representation.

[0051] Therefore, the cross-modal attention mechanism can dynamically adjust the weights between modalities according to the needs of different scenarios, avoiding the disadvantages of artificially setting fixed weights. For example: in some cases, vehicle image features (such as damage degree, maintenance status) may be more important than customer evaluations. At this time, the model will automatically increase the weight of the image modality; while analyzing customer loyalty, time-series data (such as customers' past consumption frequency) may be more critical, and the model will give higher weights to time-series features. This adaptive weighting method can ensure that the model can obtain the optimal feature fusion effect in different business scenarios, thereby improving the prediction accuracy.

[0052] S130: Input the multimodal feature group into the AI model to generate a membership acquisition strategy for the target customer group, so as to acquire members for the target auto repair shop.

[0053] Specifically, the server has obtained the customer's multimodal data (including text, images, time series data, etc.). The server performs feature fusion on this data to obtain a comprehensive multimodal feature group. The server inputs this feature group into the AI model, allowing the AI to automatically generate a member acquisition strategy based on historical data and pattern recognition. This strategy can be used to identify high-potential member customers and develop the best marketing plan to improve member conversion rates.

[0054] Furthermore, the server has completed data processing and obtained a comprehensive feature vector for each customer, which may include: Text data: customer reviews, social media interactions, customer service chat records, etc. Image data: vehicle photos uploaded by customers, repair process captured by store surveillance, etc. Time series data: customer repair frequency, consumption habits, time interval between visits to the store, etc. After the server fuses these data, it obtains a feature vector, for example: h fused =[0.75,0.35,0.9,0.4,0.8], this vector contains the comprehensive characteristics of the client. The server will fused As input, it is fed into an AI model, such as a neural network (DNN), a support vector machine (SVM), a random forest (RF), etc., to predict the probability of the customer becoming a member. In the embodiment of the present application, the AI model is a pre-trained neural network model. The role of the AI model is to identify customer consumption and behavior patterns, determine which customers are more likely to become members, and predict customer sensitivity to different membership benefits (such as discounts, free maintenance, etc.), as well as provide the best customer development strategy, such as which services to recommend to customers and when to send marketing text messages.

[0055] For example, the server formulates specific strategies based on the probability of membership conversion: Option 1 (active push): The server sends a text message to customer A: "Dear customer, you have made 3 purchases. We sincerely invite you to join the membership and enjoy a 20% discount on the first order!" Option 2 (enhance interest): The store APP pushes a message: "Members can enjoy a free car wash once this month, only for 3 days!" Option 3 (precision marketing): If customer A does not register as a member immediately, push again after 3 days: "Your maintenance is about to expire. Join the membership and enjoy a 20% discount on renewal!"

[0056] Therefore, by collecting text, image, and time-series data, a comprehensive understanding of customers and store conditions has been achieved from multiple dimensions. Feature fusion of multi-modal data can integrate the advantages of each data source, extract richer and more accurate customer and store features, and reduce the bias brought by a single data source. Inputting the fused features into an AI model to automatically generate membership customer acquisition strategies for the target customer group can quickly respond to market changes, improve the scientificity and accuracy of decision-making. The entire process from data collection, feature fusion to strategy generation is automated, greatly reducing the workload of manual statistics and improving the processing speed and customer acquisition efficiency. Based on multi-modal data and AI analysis, a customer acquisition plan can be customized for each target auto repair store, helping the store more effectively acquire high-quality customers and improve member conversion rates and customer loyalty. Therefore, it is convenient to generate accurate membership customer acquisition strategies.

[0057] In a possible implementation, the multi-modal feature group is input into the AI model to generate a membership customer acquisition strategy for the target customer group, and the following formula is specifically used for calculation: ; Where, is the membership customer acquisition strategy, which is used to represent the probability of converting customers in the target customer group into members. is the logit value generated by the neural network part, that is, the conversion tendency describing the customer's own characteristics. is the weighted term of the distance part, which is used to represent the similarity between the customer and the customer group with a high conversion tendency, and λ is the contribution strength of the adjusted distance score in the overall logit.

[0058] Specifically, the purpose of this formula is to calculate the probability of a target customer becoming a member, which is used to guide the customer acquisition strategy. First, calculate the impact of customer characteristics on member conversion. Second, calculate the similarity between the customer and the high-conversion customer group. Finally, use the sigmoid activation function to limit the result within the range of [0,1] to obtain the member conversion probability. For example, if h fused emphasizes that the customer often consumes in this store and has a good service evaluation, then its value will be relatively large, increasing the member conversion probability. For another example, if a customer's consumption behavior is highly similar to the member group (such as the same consumption frequency and service type), then the similarity value will be relatively large, thus increasing the member conversion probability.

[0059] For example, assume that the calculation result is: h fused ·W = 1.2, λ = 0.5, ; After substituting into the above formula, the probability value is approximately 0.832. That is, the probability of this customer becoming a member is approximately 83.2%, and the auto repair store can give priority to pushing membership preferential information to it.

[0060] In a possible implementation, referring to Figure 2 , Figure 2 is another process schematic diagram of an AI-based membership customer acquisition method for auto repair shops provided by an embodiment of the present application. It includes steps S210 to S230, and the above steps are as follows: S210, extracting a target probability value from the membership customer acquisition strategy; S220, comparing the target probability value with a preset probability threshold; S230, if it is determined that the preset condition indicates that the target probability value is greater than or equal to the target probability threshold, then performing membership customer acquisition on the customers in the target customer group who meet the preset conditions.

[0061] Specifically, the target probability value is an indicator that measures the possibility of a certain customer becoming a member. It is calculated by an AI model based on multimodal data (such as customer consumption records, social media evaluations, repair histories, etc.) and outputs a value between 0 and 1, representing the conversion probability of this customer. In the embodiment of the present application, a neural network is used to comprehensively evaluate customer characteristics and output a logit value. When formulating a membership customer acquisition strategy, an enterprise needs to set a preset probability threshold, that is, the conversion probability of a customer must reach a certain level to be considered a high-potential member. This threshold can be set through historical data analysis or business experience. For example: if the average conversion probability of customers who have successfully converted into members in the past is 0.7, then the preset threshold can be set to 0.7; if you want to screen more precise target customers, you can increase the threshold, such as 0.8, and vice versa, decrease it, such as 0.6.

[0062] For example, if the calculated membership probability is 0.835 and the preset threshold is 0.8, since 0.835 is greater than 0.8, the server performs membership customer acquisition, such as sending a text message: "Dear customer, you have met the VIP membership qualification. Join the membership and enjoy an 80% discount immediately!" At the same time, the APP pushes: "Congratulations! Your consumption has reached the membership standard. Click to activate member privileges immediately!" In addition, when the calculated membership probability is 0.72, since 0.7 is less than 0.8, the server does not perform membership customer acquisition for the time being, but first recommends a maintenance package to the customer to increase the consumption frequency and wait for a higher conversion signal.

[0063] Therefore, the membership customer acquisition strategy based on the target probability value is a precision marketing method. Its key processes include: extracting the target probability value of a customer to measure the possibility of their becoming a member; comparing the target probability value with the preset threshold to screen high-potential customers; performing membership customer acquisition on eligible customers and using discounts, pushes, personalized recommendations, etc. to increase the conversion rate. This strategy can improve the marketing precision, reduce resource waste, enhance customer loyalty, and thus help enterprises gain more growth space in the highly competitive market.

[0064] The present application also provides an AI-based membership customer acquisition device for auto repair shops. Referring toFigure 3 , Figure 3 This is a schematic diagram of the modules of a membership acquisition device for auto repair shops based on AI provided by an embodiment of the present application. The membership acquisition device for auto repair shops is a server, and the server includes an acquisition module 31 and a processing module 32. Among them, the acquisition module 31 acquires multimodal data for a target auto repair shop and a target customer group. The multimodal data includes text data, image data, and time series data. The text data includes customer evaluations and social media information. The image data includes vehicle condition photos and store environments. The time series data includes repair records and consumption trajectories. The processing module 32 performs feature fusion on the multimodal data to obtain a multimodal feature group. The processing module 32 inputs the multimodal feature group into an AI model to generate a membership acquisition strategy for the target customer group, so as to acquire members for the target auto repair shop.

[0065] In a possible implementation manner, the acquisition module 31 acquires multimodal data for a target auto repair shop and a target customer group, specifically including: the acquisition module 31 receives the original data sent by the store terminal device for the target auto repair shop and the target customer group. The processing module 32 performs word segmentation, cleaning, and standardization processing on the data related to the text part in the original data to obtain text data. The processing module 32 performs normalization and data augmentation on the data related to the image part in the original data to obtain image data. The processing module 32 performs smoothing, normalization, and time window segmentation on the data related to the time series part in the original data to obtain time series data. The processing module 32 generates multimodal data according to the text data, image data, and time series data.

[0066] In a possible implementation manner, the processing module 32 performs feature fusion on the multimodal data to obtain a multimodal feature group, and specifically calculates using the following formula: ; where h fused is the multimodal feature group, LayerNorm is used to normalize the fused feature vector to enhance the model stability and convergence speed, M is the number of modalities, m is the current modality, g (m) is the gating coefficient, A (m) is the attention output, is the element-wise multiplication, which is used to multiply g (m) with A (m) to achieve dynamic weighting for each dimension.

[0067] In a possible implementation manner, g (m) is specifically calculated using the following formula: ; where represents the Sigmoid activation function, which is used to map the output value between 0 and 1 as the gating factor for this modality. is the gating weight matrix, h (m) is the input feature vector representing the m-th modality. is the bias term vector.

[0068] In a possible implementation, A (m) is specifically calculated using the following formula: In a possible implementation, A (m) is specifically calculated using the following formula: ; where Q (m) is the query of modality m after linear transformation, k is all modalities, is used to calculate the attention weights between the query of each modality m and the keys of all modalities k, obtaining a probability distribution to represent how modality m affects the information of other modalities.

[0069] In a possible implementation, the multi-modal feature group is input into the AI model to generate a membership acquisition strategy for the target customer group, which is specifically calculated using the following formula: ; where is the membership acquisition strategy, which is used to represent the probability of converting customers in the target customer group into members. is the logit value generated by the neural network part, that is, the conversion tendency describing the characteristics of the customers themselves. is the weighted term of the distance part, which is used to represent the similarity between the customer and the customer group with high conversion tendency, and λ is used to adjust the contribution of the distance score in the overall logit.

[0070] In a possible implementation, the processing module 32 extracts the target probability value from the membership acquisition strategy; the processing module 32 compares the target probability value with the preset probability threshold; if the processing module 32 determines that the preset condition indicates that the target probability value is greater than or equal to the target probability threshold, it will acquire members for the customers in the target customer group who meet the preset conditions.

[0071] It should be noted that when the device provided in the above embodiments realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0072] The present application also provides an electronic device. Referring to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include: at least one processor 41, at least one network interface 44, a user interface 43, a memory 45, and at least one communication bus 42.

[0073] Among them, the communication bus 42 is used to realize the connection and communication between these components.

[0074] Among them, the user interface 43 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 43 may further include a standard wired interface and a wireless interface.

[0075] Among them, the network interface 44 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface).

[0076] Among them, the processor 41 may include one or more processing cores. The processor 41 uses various interfaces and lines to connect various parts within the entire server, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 45, as well as calling data stored in the memory 45, it executes various functions of the server and processes data. Optionally, the processor 41 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 41 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 41 and may be implemented separately by a single chip.

[0077] Among them, the memory 45 may include a Random Access Memory (RAM), or may include a Read-Only Memory. Optionally, the memory 45 includes a non-transitory computer-readable storage medium. The memory 45 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 45 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 45 may also be at least one storage device located far from the aforementioned processor 41. As Figure 4 shown, in the memory 45 as a computer storage medium, there may be included an operating system, a network communication module, a user interface module, and an application program of a method for attracting members of an auto repair shop based on AI.

[0078] In Figure 4 the electronic device shown, the user interface 43 is mainly used to provide an input interface for the user to obtain the data input by the user; and the processor 41 can be used to call the application program of a method for attracting members of an auto repair shop based on AI stored in the memory 45. When executed by one or more processors, the electronic device is caused to execute the method as described in one or more of the above embodiments.

[0079] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0080] This application also provides a computer-readable storage medium, and the computer-readable storage medium stores instructions. When executed by one or more processors, the electronic device is caused to execute the method as described in one or more of the above embodiments.

[0081] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0082] In several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0083] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0084] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0085] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks or optical discs that can store program codes.

[0086] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited by this. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will easily think of other implementation schemes of the present disclosure after considering the specification and the disclosure of the practical truth. The present application aims to cover any variations, uses or adaptive changes of the present disclosure. These variations, uses or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. An AI-based method for attracting customers for auto repair shop members, characterized in that, The method includes: Obtaining multimodal data for a target auto repair store and a target customer group, where the multimodal data includes text data, image data, and time series data. The text data includes customer evaluations and social media information. The image data includes vehicle condition photos and store environment. The time series data includes repair records and consumption trajectories; Performing feature fusion on the multimodal data to obtain a multimodal feature group; Inputting the multimodal feature group into an AI model to generate a membership customer acquisition strategy for the target customer group, so as to acquire members for the target auto repair store.

2. The AI-based method for attracting members to auto repair shops according to claim 1, wherein The obtaining of the multimodal data for the target auto repair store and the target customer group specifically includes: Receiving the original data sent by the store terminal device for the target auto repair store and the target customer group; Performing word segmentation, cleaning, and standardization processing on the data related to the text part in the original data to obtain the text data; Performing normalization and data augmentation on the data related to the image part in the original data to obtain the image data; Performing smoothing, normalization, and time window segmentation on the data related to the time series part in the original data to obtain the time series data; Generating the multimodal data based on the text data, the image data, and the time series data.

3. The AI-based membership customer acquisition method for auto repair shops according to claim 1, characterized in that The performing of the feature fusion on the multimodal data to obtain a multimodal feature group specifically uses the following formula for calculation: ; Among them, h fused is a multi-modal feature group. LayerNorm is used to normalize the fused feature vectors to enhance the model stability and convergence speed. M is the number of modalities, m is the current modality, g (m) is the gating coefficient, A (m) is the attention output, is the element-wise multiplication, which is used to multiply g (m) by A (m) to achieve dynamic weighting for each dimension.

4. The AI-based method for attracting customers for auto repair shop members according to claim 3, wherein The said g (m) Specifically, the calculation is performed using the following formula: ; Among them, represents the Sigmoid activation function, which is used to map the output value between 0 and 1 as the gating factor of this modality. is the gating weight matrix, and h (m) is the input feature vector representing the m-th modality. is the bias term vector.

5. The AI-based method for attracting members in auto repair shops according to claim 3, characterized in that, The said A (m) Specifically, it is calculated using the following formula: ; Among them, Q (m) is the query of mode m after linear transformation, and k is all modes. It is used to calculate the attention weights between the query of each mode m and the keys of all modes k, obtaining a probability distribution to represent how mode m affects the information of other modes.

6. The AI-based method for attracting members in an auto repair shop according to claim 1, wherein, The inputting of the multimodal feature group into the AI model to generate a membership customer acquisition strategy for the target customer group specifically uses the following formula for calculation: ; Among them, is a membership customer acquisition strategy, which is used to represent the probability of converting customers in the target customer group into members. is the logit value generated by the neural network part, that is, the conversion tendency describing the customer's own characteristics. is the weighted term of the distance part, which is used to represent the similarity between the customer and the customer group with high conversion tendency. λ is used to adjust the contribution of the distance score in the overall logit.

7. The AI-based method for attracting members in auto repair shops according to claim 1, characterized in that The method further includes: Extracting a target probability value from the membership customer acquisition strategy; Comparing the target probability value with a preset probability threshold; If it is determined that the preset condition indicates that the target probability value is greater than or equal to the target probability threshold, then acquiring members for the customers in the target customer group who meet the preset condition.

8. An AI-based membership customer acquisition device for auto repair shops, characterized in that, The auto repair store membership customer acquisition device includes an acquisition module (31) and a processing module (32), where The acquisition module (31) is used to obtain multimodal data for a target auto repair store and a target customer group, where the multimodal data includes text data, image data, and time series data. The text data includes customer evaluations and social media information. The image data includes vehicle condition photos and store environment. The time series data includes repair records and consumption trajectories; The processing module (32) is used to perform feature fusion on the multimodal data to obtain a multimodal feature group; The processing module (32) is further used to input the multimodal feature group into an AI model to generate a membership customer acquisition strategy for the target customer group, so as to acquire members for the target auto repair store.

9. An electronic device, characterized in that, The electronic device includes a processor (41), a memory (45), a user interface (43), and a network interface (44). The memory (45) is used to store instructions. Both the user interface (43) and the network interface (44) are used to communicate with other devices. The processor (41) is used to execute the instructions stored in the memory (45) so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-dimensional auxiliary expansion store business method and multi-agent analysis system

    CN120725719A