Intelligent Onboarding and Management System for Pet Service Merchants

CN122573229APending Publication Date: 2026-08-14SHENZHEN YUNCHUANG YOUYI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610595252.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]为了解决现有系统无法根据实际运营结果进行动态调整,导致入驻评估与后续管理脱节的技术问题,本申请提供了一种宠物服务商家智能入驻与管理系统

Benefits of technology

本申请提供了一种宠物服务商家智能入驻与管理系统,包括:画像生成模块,用于根据入驻商家在测试场景下的多模态感知数据和交互决策数据,生成入驻商家画像;运营预测模块,用于依据入驻商家在运营场景下产生的实时服务数据,对入驻商家画像进行更新,生成预测运营结果;场景优化模块,用于对比预测运营结果与未来实际运营结果间的偏差,对测试场景进行优化处理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122573229A_ABST
    Figure CN122573229A_ABST
Patent Text Reader

Abstract

This invention relates to the field of pet system technology, specifically to an intelligent onboarding and management system for pet service merchants, comprising: a profile generation module, used to generate a profile of the onboarding merchant based on multimodal perception data and interaction decision data of the onboarding merchant in a test scenario; an operation prediction module, used to update the profile of the onboarding merchant based on real-time service data generated by the onboarding merchant in an operation scenario, and generate predicted operation results; and a scenario optimization module, used to compare the deviation between the predicted operation results and the actual future operation results, and optimize the test scenario, thereby solving the technical problem that existing systems cannot dynamically adjust according to actual operation results, leading to a disconnect between onboarding assessment and subsequent management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pet system technology, specifically to a smart onboarding and management system for pet service merchants. Background Technology

[0002] Currently, pet service merchant onboarding and management systems on the market are mainly used to conduct qualification reviews and continuous management of merchants applying to join the platform in order to ensure the quality of platform services. They typically review the qualifications and historical information submitted by merchants during the onboarding stage, and generate statistical reports based on their operational data during the management stage.

[0003] However, because the review and evaluation during the onboarding phase is static and pre-set, it cannot provide dynamic feedback and self-optimization based on the actual operational results of merchants. As a result, the review and evaluation will gradually become out of touch with the actual situation over time, and it will be impossible to accurately predict and adapt to the ever-changing service capabilities and risks of merchants. Summary of the Invention

[0004] To address the technical problem that existing systems cannot dynamically adjust based on actual operational results, leading to a disconnect between onboarding assessment and subsequent management, this application provides an intelligent onboarding and management system for pet service merchants.

[0005] The intelligent onboarding and management system for pet service merchants provided in this application adopts the following technical solution: A smart onboarding and management system for pet service merchants, comprising: The profile generation module is used to generate profiles of merchants based on their multimodal perception data and interaction decision data in the test scenario. The operation prediction module is used to update the profile of merchants based on real-time service data generated by merchants in operation scenarios and generate predicted operation results. The scenario optimization module is used to compare the deviation between the predicted operational results and the actual future operational results, and to optimize the test scenarios.

[0006] Furthermore, prior to the steps based on the multimodal perception data and interaction decision data of the merchants in the test scenario, the following steps are included: Based on the merchant's application for joining the platform, a test scenario is generated, and in response to the merchant's interactive operations in the test scenario, a data collection instruction is generated. Based on the data acquisition command, after triggering the sensors and interactive devices deployed in the test scenario, the initial multimodal data of the sensors and interactive devices are collected; Feature extraction and fusion are performed on the initial multimodal data to obtain multimodal sensing data; Multimodal perception data is fed back to the test scenario to drive its evolution; Record the interactive operations of merchants as the test scenarios evolve, and obtain interactive decision data.

[0007] Furthermore, the steps of feature extraction and fusion of the initial multimodal data to obtain multimodal sensing data include: By extracting spatial visual information from consecutive video frames in the video stream and calculating motion change information between each video frame in the consecutive video frames, the spatial visual information and motion change information are superimposed to obtain the spatiotemporal features of the video. After extracting the spectrogram from the audio stream, the frequency band energy distribution and temporal envelope features in the spectrogram are identified, and the frequency band energy distribution and temporal envelope features are calculated to obtain the audio acoustic features. The video spatiotemporal features and audio acoustic features are aligned in the time dimension, and the aligned video spatiotemporal features and audio acoustic features are spliced ​​and weighted to obtain audiovisual features. After extracting numerical sequences including temperature, humidity, and object location from environmental IoT data, the correlation weights between each data point in the numerical sequence and each dimension of the audiovisual features are calculated. After recalibrating the features of each dimension based on the association weights, the recalibrated features of each dimension are concatenated with the numerical sequence to obtain multimodal sensing data.

[0008] Furthermore, based on the multimodal perception data and interaction decision data of the merchants in the test scenario, the steps to generate a merchant profile include: Spatial structure analysis and temporal change analysis are performed on multimodal sensing data to extract environmental quality indicators; Perform sequence analysis and pattern recognition on interactive decision data to extract behavioral pattern indicators; Calculate the statistical correlation between environmental quality indicators and behavioral pattern indicators, and generate an environmental-behavior correlation matrix; Based on the environment-behavior correlation matrix, environmental quality indicators and behavioral pattern indicators are weighted and fused to obtain multiple capability indices; By combining preset industry benchmark parameters, multiple capability indicators are normalized and standardized to obtain a profile of the merchants who have joined the platform.

[0009] Furthermore, based on the real-time service data generated by merchants in operational scenarios, the steps to update the merchant profiles and generate predicted operational results include: Based on the operational behavior characteristics of real-time service data, the change in operational behavior characteristics relative to the baseline characteristics in the profile of merchants is calculated to obtain the feature offset vector. Based on the feature offset vector, the various capability indicators in the profile of the merchants are weighted and adjusted to obtain the updated capability indicators. Based on updated capability indicators and real-time service data, the capability indicators for future periods are calculated to obtain predicted operational results.

[0010] Furthermore, based on the updated capability indicators and real-time service data, the steps to calculate the capability indicators for future periods and obtain the predicted operational results include: Obtain each historical capability indicator, calculate the historical change baseline rate of each historical capability indicator, and calculate the instantaneous change rate of each capability indicator based on real-time service data. By weighting and fusing the historical baseline rate of change and the instantaneous rate of change, the predicted rate of change for each capability indicator is obtained; Based on the predicted rate of change and each updated capability indicator, the predicted values ​​of each capability indicator in the future period are obtained by recursive calculation. Analyze the statistical correlation between various operational events and various update capability indicators in real-time service data, and construct an operational event-capability impact correlation matrix; Based on the operational event-capability impact correlation matrix, the predicted values ​​are corrected to obtain the predicted operational results.

[0011] Furthermore, by comparing the deviation between the predicted operational results and the actual future operational results, the steps for optimizing the test scenario include: Calculate the difference between the predicted values ​​of each first capability indicator in the predicted operational results and the actual values ​​of each second capability indicator in the actual future operational results to obtain the difference sequence. Identify the target capability indicators that exceed the preset difference in the difference sequence, map the target capability indicators back to the test scenario to obtain the dimensions to be optimized, and obtain the target operational events and the conditions for the occurrence of the target operational events that lead to the existence of the target capability indicators by retrieving future actual operational results. Based on the target operational events and their occurrence conditions, the virtual event library corresponding to the dimension to be optimized is supplemented, and the trigger probability and event difficulty of each virtual event in the virtual event library are calculated in order to optimize the test scenario.

[0012] Beneficial effects achieved: This application provides an intelligent onboarding and management system for pet service merchants, including: a profile generation module, used to generate a profile of the onboarding merchant based on the multimodal perception data and interaction decision data of the onboarding merchant in the test scenario; an operation prediction module, used to update the profile of the onboarding merchant based on the real-time service data generated by the onboarding merchant in the operation scenario and generate predicted operation results; and a scenario optimization module, used to compare the deviation between the predicted operation results and the future actual operation results and optimize the test scenario.

[0013] In this application, a closed-loop feedback optimization mechanism is constructed to address the disconnect issue. Its core lies in dynamically linking the test scenarios used for onboarding assessment with the actual operational results generated by subsequent management. First, a profile of the onboarding merchant is generated based on multimodal perception data and interaction decision data from the test scenarios. Then, during the operation phase, the merchant profile is updated based on real-time service data, generating predicted operational results. This extends the assessment of onboarding merchant capabilities from a static onboarding point to a dynamic operational process, achieving synchronized updates to management criteria. Finally, the predicted operational results are compared with future actual operational results, and the identified deviations are directly used to optimize the test scenarios. This ensures that the test scenarios relied upon for subsequent onboarding assessments of new merchants are continuously calibrated and enhanced based on feedback from historical merchant performance, allowing the onboarding assessment standards to adaptively align with actual management. Ultimately, onboarding assessment and subsequent management maintain dynamic consistency and co-evolution within a continuously iterative closed loop. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the modules of the Smart Onboarding and Management System for Pet Service Merchants in this application; Figure 2 This is a schematic diagram of the merchant onboarding selection interface of the Pet Service Merchant Intelligent Onboarding and Management System in this application. Figure 3 This is a screenshot of the interface for merchants applying to join the platform.

[0015] Explanation of icon numbers: 10. Profile generation module; 20. Operation prediction module; 30. Scene optimization module. Detailed Implementation

[0016] The following combination Figures 1 to 3 This application will be described in further detail.

[0017] This application discloses an intelligent onboarding and management system for pet service merchants.

[0018] Please refer to Figure 1 The intelligent onboarding and management system for pet service merchants proposed in this embodiment includes: The profile generation module 10 is used to generate profiles of merchants based on their multimodal perception data and interaction decision data in the test scenario; the operation prediction module 20 is used to update the profiles of merchants based on the real-time service data generated by merchants in the operation scenario and generate predicted operation results; the scenario optimization module 30 is used to compare the deviation between the predicted operation results and the actual future operation results and optimize the test scenario.

[0019] This embodiment combines the profile generation module 10, the operation prediction module 20, and the scene optimization module 30 to construct an evolutionary closed loop from evaluation and management to feedback, effectively solving the problem of the disconnect between entry evaluation and management practice.

[0020] Specifically, the profile generation module 10 is responsible for performing preliminary quantitative modeling of merchant capabilities using multimodal perception data and interaction decision data in the test scenario during the merchant onboarding stage, generating a profile of the onboarding merchant; the operation prediction module 20 links the generated profile of the onboarding merchant with real-time operation data, continuously updates the judgment of the onboarding merchant's capabilities and predicts its future performance, generating predicted operation results for future periods; and the scenario optimization module 30 compares the deviation between the predicted operation results and the actual future operation results, feeding back the actual effect of management practices to the test scenario, and optimizing the relevant data of the generated test scenario. This allows the test scenarios faced by subsequent new merchant onboarding evaluations to be continuously calibrated and enhanced based on the real operation performance of all merchants in history.

[0021] The combined effect of the above three modules transforms the system from a one-way, open-loop process into a cycle where evaluation is based on management practice feedback and management relies on dynamic evaluation results. This allows the evaluation standards for merchants to be automatically iterated and optimized in response to new problems and patterns discovered in management practice. As a result, the test scenarios used for merchant screening are always highly consistent with the core capabilities required by the platform's actual operating environment, greatly improving the predictability of the evaluation, the accuracy of management, and the adaptability of the entire system.

[0022] In one feasible implementation, steps S01 to S05 are included before generating the profile of the merchants: Step S01: Based on the merchant's application uploaded by the merchant, generate a test scenario and then generate a data collection instruction in response to the merchant's interactive operation in the test scenario.

[0023] First, regarding the steps for generating test scenarios, refer to... Figure 2 As shown, Figure 2 This is the application interface corresponding to the pet service merchant intelligent onboarding and management system of this application, which includes "Brand Merchant Onboarding", "Store Merchant Onboarding" and "Fleet Onboarding". This embodiment takes "Store Merchant Onboarding" as an example. After clicking this option, the onboarding merchant will enter... Figure 3The "Merchant Application" interface, as shown, requires users to fill in information such as "Merchant Name," "User Name," "Contact Number," "Merchant Category," "Merchant Group," "Merchant Type," and "Keywords," along with images of their operating license and relevant industry qualifications. The Pet Service Merchant Intelligent Onboarding and Management System then generates a simulated test scenario based on the submitted application information. This virtual assessment environment simulates the real-world operating environment and stress levels of a pet service business for the merchant. Specifically, it analyzes the structured data and image assets submitted by the merchant in the "Merchant Application," extracting key fields such as "Merchant Category," "Merchant Type," and images of the operating location. Subsequently, the virtual environment construction module is activated. Using a 3D modeling engine and computer vision technology, it identifies and reconstructs the uploaded location images, generating a basic 3D virtual operating scene. This scene is then endowed with interactive physical attributes such as mass, collision, and temperature through a virtual sensor network and a physics engine. Simultaneously, based on the identified merchant type, a series of challenge events are dynamically selected and configured from a pre-set simulated business event library, such as sudden pet illness, peak customer complaints, and sudden equipment malfunctions. These challenge events are organically injected into specific spatiotemporal nodes of the virtual environment in the form of behavioral logic scripts and trigger conditions. Finally, the system integrates interactive device interfaces and initializes multimodal data acquisition pipelines to capture in real time the operational sequences, decision paths, voice interactions, and even physiological response signals captured by wearable devices of merchants responding to these dynamic events. This collectively constitutes a test scenario, allowing merchants to perform a series of pre-set or free operations and decisions within the test scenario. The system can then collect multimodal perception and interaction decision data from merchants when dealing with simulated business pressures, providing a comprehensive and controllable evaluation basis for generating accurate merchant profiles.

[0024] By analyzing the interactive operations performed by merchants in the test scenario interface in real time, such as clicking virtual controls, dragging resource elements, or selecting specific options in the simulated decision-making process, a set of structured data collection instructions is immediately generated based on a pre-configured mapping relationship between the interactive operation and the data collection requirements. This set of data collection instructions specifies parameters such as the type of sensor to be triggered, the data format to be collected, the sampling rate, and the duration. This ensures that the initial multimodal data collected in the subsequent step S02 is directional raw data that is strongly correlated with the interactive operation in time and logic. This transforms the data collection mode from a fixed or all-time collection mode to an on-demand directional collection mode triggered in real time by the behavior of merchants. This ensures that the acquired multimodal perception data has a clear evaluation intent and contextual relevance from the source, greatly improving the data quality and the accuracy of subsequent feature fusion and analysis. This lays an efficient and low-noise data foundation for generating reliable merchant profiles.

[0025] Step S02: Based on the data acquisition command, after triggering the sensors and interactive devices deployed in the test scenario, the initial multimodal data on the sensors and interactive devices is acquired.

[0026] The system parses the sensor identifier, data format, sampling rate, and duration contained in the data acquisition command. Then, it sends synchronization control signals to the virtual sensors (such as HD cameras, directional microphones, temperature and humidity sensors, and motion capture devices) and interactive devices (such as touch screens and simulation instruments) deployed in the test scene specified in the data acquisition command, through the corresponding device control protocol (such as ONVIF protocol to control network cameras and Modbus protocol to read environmental sensors). This triggers these devices to start data acquisition according to the format (such as the encoding format and resolution of video streams), sampling rate (such as the sampling frequency of audio streams), and duration required by the data acquisition command.

[0027] During the data acquisition process, all data streams generated by virtual sensors and interactive devices—namely, video streams, audio streams, and environmental IoT data including temperature, humidity, and the coordinates of dynamic objects—are uniformly tagged with time series labels. These data are then aggregated and buffered according to a predetermined data encapsulation format, forming initial multimodal data that is strictly synchronized in time and logically directly related to the interactive operations. This achieves precision, synchronization, and contextualization in the acquisition of multi-source heterogeneous data. It ensures that the acquired video streams, audio streams, and environmental IoT data are not only aligned in the time dimension, but also that their acquisition motivation, content, and time period directly correspond to the specific interactive behaviors of participating merchants in the test scenario in the preceding steps. This provides a high-quality data foundation for subsequent feature extraction and fusion, enabling multimodal perception data to accurately characterize the simulated environmental state triggered by or in which the interactive operation takes place, greatly enhancing the accuracy and reliability of subsequent evaluations.

[0028] In this embodiment, the virtual sensor is a modal data that is automatically generated based on the interactive operations of the merchants in the test scenario, based on the device data of the corresponding actual sensor.

[0029] Step S03: Perform feature extraction and fusion on the initial multimodal data to obtain multimodal perception data.

[0030] Feature extraction and fusion are performed on the initial multimodal data to transform the initial multimodal data collected by S02 into a unified structured feature representation, namely multimodal perception data. This breaks down the barriers between different modal data and deeply integrates and correlates visual spatial and motion information, auditory acoustic characteristics, and the physical environment state sensed by the Internet of Things. This constructs a digital image that can characterize the dynamic environmental state and contextual information of the test scenario at a specific moment. For example, spatial visual and motion change information is extracted from the video stream to obtain video spatiotemporal features, and frequency band energy and envelope features are extracted from the audio stream to obtain audio acoustic features. The two are aligned and weighted in the time dimension to obtain audiovisual features. Numerical sequences are then extracted from the environmental IoT data and their correlation weights with the audiovisual features are calculated. Finally, the audiovisual features are recalibrated based on the weights and concatenated with the numerical sequences. This makes the generated multimodal perception data no longer a set of independent raw signals, but a set of quantitative indicators that directly correspond to the semantics of the scenario (such as safety status, busyness level, and operational compliance).

[0031] Optionally, step S03 further includes steps S031 to S035: Step S031: By extracting spatial visual information of continuous video frames from the video stream and calculating motion change information between each video frame in the continuous video stream, the spatial visual information and motion change information are superimposed to obtain the video spatiotemporal features.

[0032] The acquired video stream is decoded to obtain a sequence of consecutive video frames arranged in timestamp order. An edge detection algorithm (such as the Canny operator) is used to scan the areas of drastic grayscale changes corresponding to each video frame, identifying and outlining the basic contours of all objects in the image. Next, a region segmentation algorithm (such as a deep learning-based semantic segmentation model) is used to divide the image into multiple connected regions corresponding to different entities (such as pets, people, and equipment) based on pixel color, texture, and semantic consistency, and to determine the relative positional relationships between these regions. Simultaneously, a keypoint localization algorithm (such as a feature point detection network) is used to locate highly discriminative coordinate points on specific objects (such as pet joints or facial feature points). Texture features (such as those obtained through local binary pattern descriptors) and color distribution features (such as calculating the color histogram of the region) are extracted from the segmented regions. The contour information, region positional relationships, keypoint coordinates, and texture and color features are collectively encoded into a numerical vector, which represents the spatial visual information.

[0033] The calculation of motion change information between consecutive video frames is specifically achieved by comparing the brightness and chromaticity differences of corresponding pixel blocks between adjacent video frames and tracing the displacement trajectory of specific key points in consecutive frames. This allows the calculation of an optical flow field or motion vector field that describes the direction, speed, and acceleration of objects in the image.

[0034] Finally, a feature fusion method is used to concatenate or weight and sum the spatial visual information extracted from each frame with the motion change information calculated relative to the previous frame in the feature dimension. This generates a composite feature representation that simultaneously encodes the static spatial attributes of a single frame and the dynamic change attributes between frames, namely video spatiotemporal features. By integrating the originally separate video appearance information and motion information into a single feature expression, subsequent processing can not only perceive "what" and "where" objects are in the test scene, but also understand "how they move" and "how they change." This provides a crucial visual understanding foundation for comprehensively and accurately assessing the dynamic evolution of the environmental state caused by the merchant's operations.

[0035] Step S032: After extracting the spectrogram from the audio stream, identify the frequency band energy distribution and temporal envelope features in the spectrogram, and calculate the frequency band energy distribution and temporal envelope features to obtain the audio acoustic features.

[0036] Frame segmentation is achieved by dividing a continuous audio stream into a series of short audio stream segments of fixed length. This approximates the non-stationary audio stream as a series of stationary audio streams within short time intervals. During framing, an overlap region (e.g., a frame shift of 10 milliseconds) is set between adjacent frames to maintain the continuity of the audio stream between frames and reduce information loss. Next, a window function (such as a Hamming window) is applied to each frame of the time-domain audio stream for windowing processing. This window function is multiplied by the audio stream within the frame to reduce the spectral leakage effect caused by audio stream truncation, allowing the ends of the frame to smoothly decay to zero, thereby improving the accuracy of subsequent spectral analysis. Then, a Fast Fourier Transform (FFT) algorithm is applied to each windowed frame of the time-domain audio stream. The FFT algorithm uses an efficient butterfly computation structure to convert the discrete time-domain audio stream into a discrete frequency-domain representation, calculating the complex amplitude and phase of the audio stream at each frequency component. The output result is the spectrum. The spectra of multiple consecutive frames are arranged in chronological order and converted into a Mel scale to generate a spectrogram with time as the horizontal axis, frequency as the vertical axis, and color depth representing energy intensity. This spectrogram is a two-dimensional matrix, with rows corresponding to discretized frequency units and columns corresponding to consecutive time frames.

[0037] To identify the frequency band energy distribution in the spectrogram, the entire frequency range is divided into multiple preset key frequency bands (e.g., specific frequency intervals corresponding to human voices, pet barks, and environmental noise). Based on the logarithmic energy values ​​of specific frequencies represented by the matrix elements of the generated spectrogram at specific time frames, and according to a preset frequency band division scheme (e.g., using a Mel-scale filter bank to divide the nonlinear frequency range perceived by human hearing into multiple overlapping triangular frequency bands), a filter weight covering a specific frequency interval is defined for each target frequency band. For a given time window (usually covering multiple consecutive time frames in the spectrogram), energy integration is performed for each frequency band. That is, within the time window, the logarithmic energy values ​​of all frequency units in the spectrogram falling within the frequency range of that frequency band are multiplied by their corresponding filter weights and then summed along the frequency dimension to obtain the cumulative energy value of that frequency band within that time window. This calculation is repeated for all preset frequency bands to obtain a numerical vector, where each element corresponds to the integrated energy of a frequency band. This numerical vector describes the distribution of the audio stream's energy in different frequency intervals within a given time window, i.e., the frequency band energy distribution.

[0038] Simultaneously, the time-domain envelope features are calculated from the time-domain waveform of the audio stream. This typically involves calculating the short-time energy curve of the signal and extracting parameters such as rise time, fall time, steady-state duration, and peak-to-trough ratio from the short-time energy curve. Specifically, the audio stream undergoes the same framing and windowing process as when generating the spectrogram, resulting in a series of short-time signal frames. Next, the short-time energy of each frame is calculated, i.e., the amplitudes of all audio stream points within that frame are squared and summed (or the absolute values ​​are summed), resulting in a discrete short-time energy curve that varies with time. This curve reflects the macroscopic fluctuation profile of the audio stream amplitude, i.e., the time-domain envelope. Subsequently, each significant energy event (such as a...) is located on this curve. (A sound or a segment of speech) is detected by setting a dynamic threshold to identify the starting point (the point where energy begins to rise continuously from the background noise level), the peak point (the point where energy reaches its local maximum value), and the ending point (the point where energy falls back to the background noise level) of each energy event. Based on these key points, parameters are calculated: the rise time is defined as the time difference from the starting point to the peak point, the fall time is defined as the time difference from the peak point to the ending point, and the steady-state duration is defined as the duration during which energy remains within a certain proportion (e.g., 90% to 70%) of the peak value. The peak-to-valley ratio is usually calculated by comparing the peak value of the current energy event with the energy of the immediately preceding energy valley value (i.e., a representative value of the background noise level).

[0039] By directly concatenating the frequency band energy distribution vector and the temporal envelope feature vector in terms of dimension to form a composite feature vector, principal component analysis (PCA) is applied to reduce the dimension and decorrelate the composite feature vector. PCA calculates the eigenvalues ​​and eigenvectors of the covariance matrix that match the feature vector, and selects the eigenvector corresponding to the largest eigenvalue as the principal component direction. The composite feature vector is then projected onto the principal components, thereby reducing the feature dimension and eliminating the linear correlation between features while retaining most of the variance information. After this calculation, a numerical vector with lower dimension, more compact information, and weaker correlation between features is obtained, which is the audio acoustic feature.

[0040] Step S033: Align the video spatiotemporal features and audio acoustic features in the time dimension, and then concatenate and weight the aligned video spatiotemporal features and audio acoustic features to obtain audiovisual features.

[0041] Based on the time-series tags assigned to the video and audio streams during data acquisition, a time synchronization relationship is established between video spatiotemporal features and audio acoustic features. Since the frame rate of video spatiotemporal feature sequences is typically 30 feature vectors per second, while the frame rate of audio acoustic feature sequences is typically 100 feature vectors per second, algorithms such as linear interpolation are needed to resample the video spatiotemporal features and audio acoustic features onto the same unified time grid. First, a unified time grid is defined, covering the entire time range at fixed time intervals (e.g., 50 points per second or directly using the frame rate of one of the sequences as a reference), generating a series of target time points. Then, based on a linear interpolation algorithm, each feature is treated as a function of time and feature vector. For each target time point... If the target event point is located between two adjacent sampling points of video spatiotemporal features and audio acoustic features, then the feature vector value of the target event point is calculated by linear interpolation. That is, the feature vector of the target time point is obtained by proportionally weighting and summing the feature vectors of the two adjacent sampling points and their time difference, thereby generating an estimated feature value at the target time point. Finally, the processed video spatiotemporal feature sequence and audio acoustic feature sequence are transformed to the same target time point to ensure that there is a video spatiotemporal feature vector and an audio acoustic feature vector at each same time point, thereby completing the accurate alignment in the time dimension.

[0042] Next, the video feature vector and audio feature vector at the same time point are directly concatenated along the feature dimension to form a fused feature vector. Then, according to a preset weighting matrix, different weighting coefficients are applied to the different feature dimension components from the video stream and audio stream in the fused feature vector. These weighting coefficients reflect the differences in importance of different modal information in the current evaluation context. The weighting calculation is completed by multiplying each dimension of the vector by the corresponding coefficient. Finally, the weighted fused feature vector is output, which yields the audiovisual features.

[0043] Step S034: After extracting the numerical sequence including temperature, humidity and object location from the environmental IoT data, calculate the correlation weight between each data point in the numerical sequence and each dimension of the audiovisual features.

[0044] Read the raw data stream with timestamps reported by various sensors deployed in the test scenario; then, based on the unified time reference established in step S02, extract the data points corresponding to the audiovisual feature sequences reported by each sensor in the above raw data stream that are aligned in the time dimension, thereby forming a temperature value sequence, humidity value sequence, and value sequence describing the three-dimensional coordinates of the object that are strictly synchronized in time with the audiovisual feature sequences and arranged in chronological order.

[0045] Subsequently, the correlation weights between each data point in the numerical sequence and each dimension of the audiovisual features are calculated, specifically using the Pearson correlation coefficient algorithm: For each item in the numerical sequence and each dimension of the audiovisual feature sequence, the Pearson correlation coefficient is calculated between the value of the numerical sequence and the feature value of a certain dimension of the audiovisual feature within the same time window. The absolute value of the Pearson correlation coefficient represents the degree of linear correlation between the environmental variable and the audiovisual feature dimension in terms of their changing trends. Then, all calculated Pearson correlation coefficients are normalized (e.g., using the Softmax function) so that each audiovisual feature dimension corresponds to a set of weight coefficients from different environmental variables, and their sum is 1. This set of coefficients constitutes the correlation weights. The specific calculation process of the Pearson correlation coefficient can be found in step S13.

[0046] By quantifying the correlation weights, the dynamic correlation strength between different environmental factors (i.e., temperature and humidity changes, object movement) and different audiovisual events (such as specific actions, sounds) is characterized.

[0047] Step S035: After recalibrating the features of each dimension based on the association weights, the recalibrated features of each dimension are concatenated with the numerical sequence to obtain multimodal sensing data.

[0048] For each dimension in the audiovisual feature sequence, a comprehensive recalibration coefficient is calculated based on the association weights of that dimension with all environmental variables. This recalibration coefficient is usually obtained by weighted summation or averaging of all weights associated with that dimension.

[0049] Each feature value of the audiovisual feature sequence is multiplied by its corresponding recalibration coefficient to generate a new feature vector. Features with stronger correlation to changes in the current environmental state are enhanced, while features with weaker correlation are relatively suppressed. This process is called recalibration. Then, the recalibrated audiovisual feature sequence and the numerical sequence are directly concatenated sequentially along the feature dimensions to form a single feature vector containing all modal information. This single feature vector is the multimodal perception data.

[0050] By enhancing or suppressing audiovisual feature sequences based on quantified environmental correlation, the feature representation is made more focused on perceptual information that changes in synergy with the environmental state. Then, through cascading operations, the adjusted perceptual features are seamlessly integrated with the quantified parameters of the environment itself, thereby generating a unified multimodal perceptual data. This provides a high-quality input with highly integrated and deredundant information for subsequent scene evolution driving and merchant behavior assessment.

[0051] Step S04: Feed the multimodal perception data back to the test scenario to drive the evolution of the test scenario.

[0052] Multimodal perception data is transmitted as a real-time input to the scene engine that controls the logic and content generation of the test scenario. This scene engine has a pre-built rule base for event triggering and parameter adjustment. It parses the input multimodal perception data and identifies key state indicators, such as object position values ​​indicating that a pet has entered a high-risk area, audio acoustic features indicating the presence of loud noise, and video spatiotemporal features indicating that people are running. Based on the preset causal logic mapping relationship, it determines the evolution direction of the subsequent test scenario. For example, when a combination of high temperature and pet gathering is detected, a simulated air conditioning failure event is automatically triggered, or when a person leaves a designated area for a long time, a decision point simulating an emergency customer call is added. This allows for real-time updates of events, challenges, and parameters in the test environment, driving the entire test scenario to evolve from one state to the next more complex or targeted state.

[0053] Step S05: Record the interactive operations of merchants in the test scenario and obtain interactive decision data.

[0054] The system continuously monitors and captures all operations performed by merchants in the test scenario interface, including clicking virtual controls, dragging resources, selecting options in the decision tree, or entering commands. Whenever an operation is identified, the system accurately records its timestamp, operation type, target object, and specific parameters. This record is then bound to the current state identifier of the test scenario when the operation is triggered. As the test scenario evolves based on feedback multimodal perception data, these operation records bound to the scenario state are continuously linked in chronological order, ultimately forming a time-series data set containing a complete sequence of operations and their corresponding scenario context—that is, interactive decision data. This generates a complete set of decision-making process data that not only records "what the merchants did" but also accurately records "in what dynamic scenario state they performed this action." This provides a complete, coherent, and highly interpretable data foundation for subsequent analysis of the merchants' decision-making logic, response modes, and interactions with the dynamic environment.

[0055] In one feasible implementation, the specific execution steps of the image generation module include steps S11 to S15: Step S11: Perform spatial structure analysis and temporal change analysis on the multimodal sensing data to extract environmental quality indicators.

[0056] For spatial structure analysis, feature dimensions representing static spatial layout are separated from multimodal perception data, such as object outlines and relative position information parsed from recalibrated audiovisual feature sequences, and object position coordinates obtained from cascaded numerical sequences. Environmental safety scores are quantified by calculating whether the minimum distance between key objects, such as pets, people, and equipment, is below the safety threshold, whether hazardous materials are placed in isolated areas, and whether the physical separation of different functional areas is clear. At the same time, space utilization is calculated by statistically analyzing the area ratio and time ratio of each functional area occupied within a given time period.

[0057] For time-series change analysis, feature dimensions with time-series characteristics are extracted from multimodal perception data, such as continuous readings of temperature and humidity, trajectory sequences of object movement, and start and end markers of sound events in audio features. By analyzing the change patterns of these sequence data, such as checking whether the temperature fluctuates within the permissible range and whether the duration and sequence of cleaning and disinfection processes meet preset standards, an operational compliance score is calculated. In this way, the multimodal perception data is transformed into a set of environmental quality indicators with clear business meanings that quantify and evaluate the test scenario status from two dimensions: the rationality of spatial layout and the standardization of time-series operations. This provides quantitative environmental input for the correlation analysis with the behavior patterns of resident merchants in subsequent steps.

[0058] Step S12: Perform sequence analysis and pattern recognition on the interactive decision data to extract behavioral pattern indicators.

[0059] Sequence analysis is performed on the interactive decision data. The timestamps of all records are extracted from the interactive decision data in the order of the operation occurrence time to form an increasing timestamp sequence. The difference between adjacent timestamps in the timestamp sequence is calculated to obtain the time interval data representing the interval between consecutive operations. Then, the arithmetic mean of all time interval data is calculated to obtain the average decision delay index. At the same time, the standard deviation of these time interval data is calculated to obtain the operation rhythm fluctuation index. Furthermore, the operation subsequences that occur frequently and have a fixed order are analyzed. For example, after simulating a pet's discomfort, the sequence of three operations—checking vital signs, reviewing files, and starting ventilation—occurs consecutively to identify its routine operation process.

[0060] Next, pattern recognition is performed, and the usual operating procedures of the merchants are dynamically warped with multiple preset ideal response modes (such as standard emergency procedures, efficient service procedures, etc., which are defined as standardized operation sequences in specific scenario states). Specifically, a distance matrix is ​​constructed to calculate the minimum cumulative distance between the usual operating procedures and each preset ideal response mode. The distance metric is usually based on the semantic similarity between operation types or a preset cost function (e.g., the distance is 0 for the same operation and 1 for different operations). Non-linear alignment on the time axis is allowed to handle the differences in operation execution speed, thereby finding the optimal curved path. Based on this, the similarity score between the usual operating procedures and each preset ideal response mode is obtained through normalization (e.g., mapping the distance to the 0-1 interval, where 1 represents a complete match). At the same time, deviation paths are extracted from the curved paths generated by dynamic time warping. These deviation paths show the matching status of corresponding operations in the usual operating procedures and preset ideal response modes, as well as the positions of insertion, deletion, or replacement operations. Next, an operational sequence compliance indicator is extracted. This indicator is directly defined by the similarity score between the usual operating procedure and the corresponding preset ideal response mode. A higher similarity score indicates better compliance. Simultaneously, a process optimization indicator is extracted. By analyzing deviation paths, if the usual operating procedure requires fewer steps or takes less time to achieve the same goal compared to the preset ideal response mode (i.e., the deviation path shows some operations skipped or merged), the optimization level is high. Conversely, if the deviation path shows additional redundant operations, the optimization level is low. The optimization level can be quantified as a score based on path length and operational efficiency. Finally, by comparing the matching degree between the usual operating procedure and preset ideal response modes with different tendencies (e.g., risk-averse modes typically include more check steps and conservative operations, while efficiency-first modes have simpler steps), if the matching degree between the usual operating procedure and the risk-averse mode is significantly higher than that of the efficiency-first mode, it is determined to be risk-averse; otherwise, it is efficiency-first. This yields the decision-making tendency mode.

[0061] Finally, the average decision delay, operational rhythm volatility, operational sequence compliance, process optimization degree, and decision tendency pattern calculated from sequence analysis and pattern recognition are quantified into specific numerical values. Specifically: for the average decision delay, its value is directly represented by the arithmetic mean of all adjacent operational time intervals; for operational rhythm volatility, its value is represented by the standard deviation of the aforementioned time interval sequence; for operational sequence compliance, its value is represented by the similarity score between the habitual operational process calculated by the dynamic time warping algorithm and the most relevant preset ideal response pattern. This similarity score is mapped to the interval between 0 and 1 using the formula (1 - minimum cumulative distance / maximum possible distance), where 1 represents complete compliance; for process optimization degree, its value is obtained by analyzing the deviation path derived from dynamic time warping. Quantification is performed, for example, by calculating the ratio of the number of effective operation steps required to achieve the same target state as the preset ideal response mode using the usual operating procedures to the number of steps in the ideal mode, thus obtaining an optimization score. For the decision-making tendency mode, its value is quantified by comparing the matching degree scores of the usual operating procedures with the risk-averse ideal mode and the efficiency-first ideal mode. Specifically, the difference between the risk-averse matching degree score and the efficiency-first matching degree score can be calculated, or the ratio of the risk-averse matching degree score to the sum of the two can be calculated, thus obtaining a continuous value. A positive value or a high ratio indicates a tendency towards risk-averse behavior, while a negative value or a low ratio indicates a tendency towards efficiency-first behavior, thus forming a behavioral pattern indicator. This provides data input on the subjective behavior of resident merchants for subsequent correlation analysis and comprehensive capability assessment with environmental quality indicators.

[0062] Step S13: Calculate the statistical correlation between environmental quality indicators and behavioral pattern indicators to generate an environmental-behavior correlation matrix.

[0063] It should be noted that during the data acquisition phase, a unified high-precision time source is used to assign a millisecond-level timestamp with a unified benchmark to each multimodal sensing data and each interactive decision data. During the data processing phase, a fixed time window length is defined, and then the entire test time axis is divided into continuous non-overlapping or partially overlapping sections using this time window length as the sliding unit. For each time window, all multimodal sensing data falling within the time window are filtered based on the timestamp, and the environmental quality index of the time window is calculated in real time. At the same time, all interactive decision data recorded within the time window are filtered based on the timestamp, and the behavioral pattern index of the time window is calculated, thus forming time-aligned environmental quality index and behavioral pattern index.

[0064] For each pair of environmental quality indicators and behavioral pattern indicators, the ratio of the covariance to the standard deviation of each indicator is calculated using the Pearson correlation coefficient formula to obtain a correlation coefficient ranging from -1 to 1. First, assuming that after time alignment, a certain environmental quality indicator X and a certain behavioral pattern indicator Y each have n values ​​observed within the same time window, forming separate sequences... and Next, the covariance of the two sequences is calculated using the following formula: ,in and These are the sample means of sequences X and Y, respectively, and the sample standard deviations of the two sequences are calculated using the following formulas: , Finally, the correlation coefficient r is calculated using the formula... , where r=1 indicates perfect positive correlation, r=-1 indicates perfect negative correlation, and r=0 indicates no linear correlation.

[0065] After calculating the correlation coefficients between all environmental quality indicators and behavioral pattern indicators, these correlation coefficients are organized into a matrix, where the rows of the matrix correspond to each environmental quality indicator, the columns correspond to each behavioral pattern indicator, and each element in the matrix is ​​the corresponding correlation coefficient, thereby generating an environment-behavior correlation matrix to reveal the degree of correlation between different environmental dimensions and different behavioral pattern dimensions.

[0066] The environment-behavior correlation matrix is ​​a mathematical matrix in which rows represent environmental quality indicators and columns represent behavioral pattern indicators. Each element in the matrix is ​​the Pearson correlation coefficient between the corresponding indicator pairs, which is used to quantify the statistical correlation strength between environmental conditions and business behavior patterns.

[0067] Step S14: Based on the environment-behavior correlation matrix, the environmental quality indicators and behavioral pattern indicators are weighted and fused to obtain multiple capability indices.

[0068] The capability index includes the risk response capability index, the service stability index, and the resource scheduling efficiency index.

[0069] Based on the preset capability definition mapping rules, specific correlation coefficients related to the target capability are extracted from the environment-behavior correlation matrix as weight coefficients. For example, when calculating the risk response capability index, the correlation coefficients between the behavioral pattern indicators "operational sequence compliance" and "decision tendency pattern" and the environmental quality indicators "environmental safety score" and "operational compliance score" are extracted. These correlation coefficients are then normalized to obtain a set of weight coefficients for weighting. These coefficients are then used to perform a weighted summation calculation on the corresponding environmental quality indicators and behavioral pattern indicators. That is, the risk response capability index = ∑(weight coefficient * corresponding value).

[0070] The calculation of the service stability index and the resource scheduling efficiency index follows the same logic, but the sources of the weighting coefficients and the corresponding indicators are different. For example, the service stability index focuses more on the correlation coefficients between "operation rhythm fluctuations", "process optimization", "space utilization", and "operational compliance score", while the resource scheduling efficiency index focuses on the correlation coefficients between "process optimization", "average decision delay", "space utilization", and "operational compliance score". This makes each capability index not a simple score of the performance of the environment or behavior, but a quantitative result that deeply integrates the interactive relationship of "what behavior mode adopted under what environmental conditions" of the merchants in the test. This makes the capability assessment more in line with the complex situational dependence in real operation and significantly improves the accuracy and interpretability of the assessment.

[0071] Step S15: Combine preset industry benchmark parameters to normalize and standardize multiple capability indicators to obtain a profile of the merchants.

[0072] It should be noted that the preset industry benchmark parameters are key distribution parameters derived from a large amount of historical test data from the same industry or scenario. Specifically, these include the industry average, industry standard deviation, industry highest value, and industry lowest value for each capability index.

[0073] A Z-score-based standardization method combined with linear scaling is employed for mapping. For example, for a merchant's capability index, the difference between it and the industry average is first calculated, then divided by the industry standard deviation to obtain a standard score. This aims to eliminate dimensions and reflect the degree to which the merchant deviates from the industry average. Then, to obtain a more intuitive and consistent score, the standard score is mapped to a preset fixed range (e.g., 0 to 100) using a linear function, resulting in a capability index score: capability index score = industry average + 10 * standard score. Finally, all capability index scores after the above standardization and normalization processes are integrated to form a quantitative merchant profile. By converting the individual merchant's capability index into a capability index score relative to the overall industry level, the comparison barriers caused by different dimensions and numerical ranges between different capability dimensions are eliminated. This allows the generated merchant profile to not only clearly depict the merchant's strengths and weaknesses in each dimension but also accurately position the merchant in the competitive industry environment, thus providing an objective and fair basis for the platform's differentiated resource matching, service recommendations, or risk level classification.

[0074] In one feasible implementation, the specific execution steps of the operation forecasting module include steps S21 to S23: Step S21: Based on the operational behavior characteristics of real-time service data, calculate the change in operational behavior characteristics relative to the baseline characteristics in the profile of merchants, and obtain the feature offset vector.

[0075] A set of quantifiable operational behavior features is extracted from real-time service data according to preset rules. These features include order fulfillment efficiency, sentiment scores of customer interaction content, and completion scores of key service process nodes. The definitions of these operational behavior features are aligned with the dimensions used to generate merchant profiles. Simultaneously, capability index scores that can serve as comparison benchmarks are retrieved from stored merchant profiles. These include risk response capability index, service stability index, and resource scheduling efficiency index. Then, for each pair of corresponding operational behavior features and capability index scores, the absolute difference is calculated. This involves subtracting the corresponding capability index score from the current operational behavior feature value to obtain a numerical value representing the change in that dimension.

[0076] Finally, the change values ​​of all dimensions are arranged in a predefined order to form a feature offset vector, which describes the direction and magnitude of the deviation of the various operational performances of the merchants from their baseline capability level. This provides data input for the subsequent dynamic adjustment and updating of the capability profile of the merchants, realizing the transformation of the merchant's operational status from static snapshot evaluation to continuous dynamic tracking.

[0077] Among them, the order fulfillment efficiency is aligned with the service stability index, the sentiment score is aligned with the risk response capability index, and the completion score is aligned with the resource scheduling efficiency index.

[0078] Step S22: Based on the feature offset vector, the various capability indicators in the profile of the merchants are weighted and adjusted to obtain the updated capability indicators.

[0079] An adjustment weight coefficient is set for the change value of each dimension in the feature offset vector and the capability index score corresponding to the merchant profile. This adjustment weight coefficient can be pre-configured based on the reliability of the historical changes of that dimension or the business importance. The calculation for the weighted adjustment of each capability index score is: Update capability index = capability index score + (corresponding change value * adjustment weight coefficient). After performing this calculation for all capability index scores in sequence, a set of updated capability indices reflecting the impact of the latest operational status can be obtained.

[0080] Step S23: Based on the updated capability indicators and real-time service data, calculate the capability indicators for future time periods to obtain the predicted operational results.

[0081] This step calculates various capability indicators for future periods based on updated capability metrics and real-time service data. Building upon the dynamic updates to the current capability status of participating merchants, it further extrapolates their future operational performance to support the platform's proactive management decisions. By integrating updated capability indicators and real-time service data, it derives quantitative values ​​for each capability indicator for future periods, thus expanding the management perspective from "what is the current state" to "what might the future state be like." This allows the system to anticipate potential enhancements, declines, or fluctuations in participating merchants' service capabilities, providing crucial, data-driven decision-making support for targeted resource allocation, risk warnings, service recommendations, or contract adjustments. Ultimately, this upgrades the operational management model from passive response to proactive intervention.

[0082] Optionally, step S23 may further include steps S231 to S235: Step S231: Obtain each historical capability indicator, calculate the historical change baseline rate of each historical capability indicator, and calculate the instantaneous change rate of each capability indicator based on real-time service data.

[0083] The system retrieves time-series data of a specific historical capability indicator from stored merchant operational records for multiple consecutive periods (e.g., the past 30 days). These periods are coded chronologically as independent variable A (e.g., day 1, day 2, up to day 30, corresponding to A values ​​of 1, 2, ..., 30). The historical capability indicator value for each period is used as the dependent variable B, forming a set of data points (A_i, B_i). A least-squares linear regression is performed, and the average slope m is calculated using the formula, where... Then, the calculated average slope is divided by the initial or average value of the historical capability indicator over multiple periods to obtain the historical change benchmark rate representing the corresponding historical capability indicator.

[0084] Simultaneously, based on real-time service data, operational behavior features aligned with capability indicators are extracted from the real-time service data of the latest time window. Following the logic of step S21, the operational behavior features are compared with the corresponding dimensions in the merchant profile to obtain a feature offset vector. Following the logic of step S22, this feature offset vector is used as an adjustment weight to adjust the capability indicator scores of each capability indicator in the merchant profile. After obtaining the updated capability indicators, the capability indicator scores stored in the previous time window are obtained. The difference between the updated capability indicators of the latest time window and the updated capability indicators of the previous time window is calculated as the change value. This change value is divided by the updated capability indicators of the previous time window to obtain the instantaneous change rate of each capability indicator.

[0085] Step S232: Weighted fusion of historical change benchmark rate and instantaneous change rate to obtain the predicted change rate of each capability indicator.

[0086] Each capability indicator is preset with a trend weight coefficient and a volatility weight coefficient. The sum of these two coefficients is 1. Their specific values ​​can be configured based on business experience or dynamically adjusted through variance analysis of historical data of the capability indicator to reflect the degree of trust in long-term patterns and recent signals.

[0087] A weighted fusion calculation is performed on each capability indicator. The calculation formula is: Predicted rate of change = (Historical benchmark rate of change * Trend weight coefficient) + (Instantaneous rate of change * Fluctuation weight coefficient). Through configurable weight coefficients, the historical benchmark rate of change, which represents the long-term stable evolution pattern, is combined with the instantaneous rate of change, which captures the impact of recent sudden events. This ensures that the obtained predicted rate of change is neither overly affected by short-term noise and becomes unstable, nor is it lagging behind due to ignoring the latest dynamics. This lays the foundation for the next step of recursive prediction.

[0088] Step S233: Based on the predicted rate of change and each updated capability indicator, perform recursive calculations to obtain the predicted values ​​of each capability indicator for the future time period.

[0089] Using each updated capability index 'a' as the initial value for recursive calculation, and the predicted rate of change 'b' as the rate of change within each recursive step, a linear recursive formula is used for calculation. The formula for calculating the predicted value Y of a certain capability index in the future time period 't' (t=1,2,3,...) is as follows: This formula is based on the assumption that the corresponding update capability index grows compounded at the predicted rate of change within a unit period.

[0090] This calculation is performed sequentially on the risk response capability index, service stability index, and resource scheduling efficiency index, and then repeated according to the number of future periods to obtain a set of predicted values ​​for each capability index in the future period.

[0091] Step S234: Analyze the statistical correlation between each operational event and each update capability indicator in the real-time service data, and construct an operational event-capability impact correlation matrix.

[0092] Discrete operational events, such as promotional activities, customer complaint handling, equipment malfunction reports, and new service launches, are identified and categorized from real-time service data. Each operational event is labeled with its occurrence time, event type, and basic attributes. Simultaneously, the update capability indicators corresponding to these operational events in time are obtained. Next, an analysis time window centered on the event occurrence point is defined for each operational event (i.e., from time T1 before the operational event to time T2 after the event). Within this time window, the average change in the update capability indicator values ​​before and after the corresponding operational event is calculated relative to the scores of the capability indicators before the event. Then, the correlation strength is quantified using the Pearson correlation coefficient algorithm. The correlation coefficient between the occurrence intensity (or frequency) sequence of each type of operational event and the change value sequence of each capability indicator is calculated. After significance testing, this correlation coefficient is used to characterize the net influence strength and direction of the corresponding operational event on that capability indicator (where positive correlation indicates promotion, and negative correlation indicates inhibition).

[0093] Finally, the event types of all operational events are used as rows of the matrix, all capability indicators are used as columns of the matrix, and the calculated correlation coefficients are filled into the corresponding positions to construct the operational event-capability impact correlation matrix.

[0094] The significance test is calculated as follows: the correlation coefficient r is calculated based on the sample data, the sample size z is determined, and then the test statistic is constructed. Its calculation formula is This statistic follows a set of degrees of freedom assuming the null hypothesis (i.e., the population correlation coefficient is 0, indicating no linear correlation between variables) holds true. of Distribution, based on calculations Values ​​and degrees of freedom By consulting The cumulative probability table yields probability values, which represent the probability of observing the correlation coefficient r of the current sample under the condition that the null hypothesis is true. For example, assuming that the correlation coefficient r = 0.3 between a certain operational event and a certain capability indicator is obtained from real-time service data analysis, and the sample size z = 30, then the following calculation is performed: The value is approximately 1.64, representing the degrees of freedom. It is 28, passed The probability value obtained from the distribution calculation is approximately 0.112. If the significance level is set to 0.05, the null hypothesis cannot be rejected because the probability value is greater than the significance level, indicating that the correlation coefficient has not passed the significance test.

[0095] in, A cumulative probability table is a pre-calculated statistical tool that lists the cumulative probability distributions under different degrees of freedom. The probability value corresponding to measuring a specific value, used in a given context. Values ​​and degrees of freedom Quickly find the corresponding probability value; the operational event-capability impact correlation matrix is ​​a mathematical matrix in which rows represent various operational events (such as promotions and complaints) and columns represent various capability indicators. The value of each element in the matrix is ​​the correlation coefficient between the corresponding event and the capability indicator, which is used to quantify the impact of different operational events on the merchant's capabilities.

[0096] Step S235: Based on the operational event-capability impact correlation matrix, the predicted values ​​are corrected to obtain the predicted operational results.

[0097] After obtaining the event types and probability values ​​of operational events, the operational event-capability impact correlation matrix is ​​queried to obtain the correlation coefficients of each event type to each capability indicator. The predicted values ​​of each capability indicator at each future time point are then corrected to obtain the capability prediction value. The correction formula is: Capability prediction value = Predicted value + ∑(Correlation coefficient * Probability value * Time decay weight). The summation iterates through all operational events that are planned to be predicted at or near the current time point. The time decay weight is used to simulate the decay effect of the corresponding operational event's impact over time.

[0098] The predicted operational results are generated by integrating all capability indicators at all future points in time.

[0099] In one feasible implementation, the specific execution steps of the scene optimization module include steps S31 to S33: Step S31: Calculate the difference between the predicted values ​​of each first capability indicator in the predicted operational results and the actual values ​​of each second capability indicator in the future actual operational results, and obtain the difference sequence.

[0100] Based on a unified time period division, future time periods are divided into equal intervals, such as dividing the next week into continuous time periods based on days, and ensuring that the data of predicted operating results and future actual operating results are aligned with this time period.

[0101] For each identical time period, the predicted values ​​of each first capability indicator (i.e., risk response capability, service stability, and resource scheduling efficiency) in the predicted operational results for that period are selected, and the actual values ​​of each second capability indicator in the corresponding period in the future actual operational results are selected. In this embodiment, the first capability indicator and the second capability indicator are the values ​​of the same set of capability dimensions from different sources.

[0102] For each pair of corresponding capability indicators, calculate the difference: difference = predicted capability value - actual capability value. This difference reflects the deviation of the actual performance of the corresponding capability indicator from the prediction within the period (positive value is underestimation, negative value is overestimation). Repeat this calculation for all capability indicators in all time periods, and arrange all the differences obtained in chronological order and capability indicator dimension order to form a difference sequence.

[0103] Step S32: Identify the target capability indicators that exceed the preset difference in the difference sequence, and map the target capability indicators back to the test scenario to obtain the dimension to be optimized. Also, by retrieving future actual operation results, obtain the target operation events that lead to the existence of the target capability indicators and the conditions for the occurrence of the target operation events.

[0104] The absolute value of each difference in the difference sequence is compared with a preset difference. When the absolute value of a difference exceeds the preset difference, the corresponding capability indicator is determined as the target capability indicator.

[0105] Next, a mapping table describing the correspondence between each capability indicator and the initial dimensions in the test scenario is queried. The identified target capability indicators are mapped back to the initial dimensions on which the test scenario was built. For example, risk response capability is mapped back to the emergency decision-making dimension. This mapping result is the dimension to be optimized.

[0106] Next, by retrieving the database of future actual operational results, the time period in which the target capability indicator deviates significantly is located, and all recorded operational events are extracted from the operational logs of that time period. Then, by tracing the event-consequence causal chain, the core events most likely to cause the actual performance of the capability indicator to deviate from the prediction and their specific background parameters (such as customer traffic, resource idle rate, time point, etc.) are identified. This information constitutes the target operational events and their occurrence conditions for the target capability indicator.

[0107] Among them, the pre-set deviation value is a critical value used to judge whether the prediction deviation of the capability indicator is significant, while the event-consequence causal chain tracing is to locate the specific event that caused the deviation by analyzing the causal relationship between operational events and the actual value of the capability indicator.

[0108] Step S33: Based on the target operational events and occurrence conditions, after supplementing the virtual event library corresponding to the dimension to be optimized, calculate the trigger probability and event difficulty of each virtual event in the virtual event library in order to optimize the test scenario.

[0109] The target operational events and their occurrence conditions are structurally analyzed to extract key elements such as event type, concurrency pressure, and time context. A virtual event template is then generated based on this, and added to the virtual event library corresponding to the dimension to be optimized. The trigger probability of each virtual event in the virtual event library is calculated. The specific calculation is based on the statistical frequency of this type of event in historical actual operational data, and comprehensively considers the probability of each prerequisite state in the test scenario. The Bayesian theorem is used for fusion calculation to obtain an updated probability value.

[0110] Simultaneously, the difficulty of each virtual event is calculated based on the actual performance of participating merchants in handling similar events in real historical data (such as average resolution time, resource consumption, and final customer satisfaction). By comparing the actual performance with a preset ideal benchmark, a difficulty coefficient is quantified. Finally, using the updated virtual event library, updated probability values, and difficulty coefficient, the logic engine and parameter generator of the test scenario are optimized. Specifically, this includes adjusting the random occurrence weight of various virtual events in the test scenario based on the probability values, and setting the resource constraint strength, time pressure parameters, and weight of the virtual event in the subsequent capability evaluation score based on the difficulty coefficient. This allows the test scenario to dynamically absorb and internalize the complexity and statistical regularities of the real world, thereby driving its continuous self-evolution. This ensures that the evaluation of future participating merchants can continuously approach and cover the core risks and efficiency bottlenecks in real operations, forming an intelligent and adaptive evaluation environment based on a closed loop of practical feedback.

[0111] The virtual event library is a set of events pre-set in the test scenario to simulate various operational challenges (such as sudden illness of pets or customer complaints). The system will add new virtual event templates and adjust trigger parameters based on historical operational feedback to continuously optimize the evaluation environment.

[0112] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A smart onboarding and management system for pet service merchants, characterized in that, include: The profile generation module is used to generate profiles of merchants based on their multimodal perception data and interaction decision data in the test scenario. The operation prediction module is used to update the profile of the merchants based on the real-time service data generated by the merchants in the operation scenario and generate predicted operation results. The scenario optimization module is used to compare the deviation between the predicted operational results and the actual future operational results, and to optimize the test scenario.

2. The intelligent onboarding and management system for pet service merchants according to claim 1, characterized in that, The step prior to the step of using multimodal perception data and interaction decision data from merchants in the test scenario includes: Based on the merchant onboarding application uploaded by the onboarding merchant, after generating the test scenario, in response to the onboarding merchant's interactive operation in the test scenario, a data collection instruction is generated; According to the data acquisition command, after triggering the sensors and interactive devices deployed in the test scenario, the initial multimodal data on the sensors and interactive devices are collected; Feature extraction and fusion are performed on the initial multimodal data to obtain the multimodal sensing data; The multimodal sensing data is fed back to the test scenario to drive the evolution of the test scenario; Record the interactive operations of the merchants as the test scenario evolves, and obtain the interactive decision data.

3. The intelligent onboarding and management system for pet service merchants according to claim 2, characterized in that, The initial multimodal data includes video streams, audio streams, and environmental IoT data. The step of extracting and fusing features from the initial multimodal data to obtain the multimodal sensing data includes: By extracting spatial visual information of consecutive video frames from the video stream and calculating motion change information between each video frame in the consecutive video frames, the spatial visual information and the motion change information are superimposed to obtain the video spatiotemporal features. After extracting the spectrogram from the audio stream, the frequency band energy distribution and temporal envelope features in the spectrogram are identified, and the frequency band energy distribution and temporal envelope features are calculated to obtain the audio acoustic features; Align the video spatiotemporal features and the audio acoustic features in the time dimension, and then concatenate and weight the aligned video spatiotemporal features and audio acoustic features to obtain audiovisual features; After extracting numerical sequences including temperature, humidity, and object location from the environmental IoT data, the correlation weights between each data point in the numerical sequence and each dimension of the audiovisual features are calculated. After recalibrating the features of each dimension based on the association weights, the recalibrated features of each dimension are concatenated with the numerical sequence to obtain the multimodal sensing data.

4. The intelligent onboarding and management system for pet service merchants according to claim 1, characterized in that, The step of generating a profile of a merchant based on the multimodal perception data and interaction decision data of the merchant in the test scenario includes: Spatial structure analysis and temporal change analysis are performed on the multimodal sensing data to extract environmental quality indicators; Sequence analysis and pattern recognition are performed on the interactive decision data to extract behavioral pattern indicators; Calculate the statistical correlation between the environmental quality indicators and the behavioral pattern indicators to generate an environment-behavior correlation matrix; Based on the environment-behavior correlation matrix, the environmental quality indicators and the behavioral pattern indicators are weighted and fused to obtain multiple capability indices; By combining preset industry benchmark parameters, the multiple capability indicators are normalized and standardized to obtain the profile of the merchants who have joined the platform.

5. The intelligent onboarding and management system for pet service merchants according to claim 1, characterized in that, The step of updating the profile of the merchants based on the real-time service data generated by the merchants in the operational scenario and generating predicted operational results includes: Based on the operational behavior characteristics of the real-time service data, the change of the operational behavior characteristics relative to the baseline characteristics in the profile of the merchants is calculated to obtain the feature offset vector; Based on the feature offset vector, the capability indicators in the profile of the merchants are weighted and adjusted to obtain the updated capability indicators. Based on the updated capability indicators and the real-time service data, the capability indicators for future time periods are calculated to obtain the predicted operational results.

6. The intelligent onboarding and management system for pet service merchants according to claim 5, characterized in that, The step of calculating the predicted operational results based on the updated capability indicators and the real-time service data for future time periods includes: Obtain each historical capability indicator, calculate the historical change baseline rate of each historical capability indicator, and calculate the instantaneous change rate of each capability indicator based on the real-time service data; The historical change benchmark rate and the instantaneous change rate are weighted and fused to obtain the predicted change rate of each capability indicator; Based on the predicted rate of change and the updated capability indicators, the predicted values ​​of each capability indicator in the future time period are obtained by recursive calculation. Analyze the statistical correlation between each operational event and each updated capability indicator in the real-time service data, and construct an operational event-capability impact correlation matrix; Based on the operational event-capability impact correlation matrix, the predicted values ​​are corrected to obtain the predicted operational results.

7. The intelligent onboarding and management system for pet service merchants according to claim 1, characterized in that, The step of comparing the deviation between the predicted operational results and the actual future operational results, and optimizing the test scenario, includes: Calculate the difference between the predicted capability values ​​of each first capability indicator in the predicted operational results and the actual capability values ​​of each second capability indicator in the future actual operational results to obtain a difference sequence; Identify the target capability indicators that exceed the preset difference in the difference sequence, and map the target capability indicators back to the test scenario to obtain the dimension to be optimized. Also, by retrieving the future actual operation results, obtain the target operation events that lead to the existence of the target capability indicators and the occurrence conditions of the target operation events. Based on the target operational event and the occurrence conditions, the virtual event library corresponding to the dimension to be optimized is supplemented, and the trigger probability and event difficulty of each virtual event in the virtual event library are calculated in order to perform the optimization processing on the test scenario.