Multi-source business data fusion analysis and visualization control system for debt collection process
Patent Information
- Application Number
- CN202611078853.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-01
AI Technical Summary
通话录音转写文本仅用于话术合规检查,还款操作日志仅用于账务核对,外访轨迹仅用于考勤管理,各数据源之间缺乏有效的融合机制,无法形成对催收过程协同状态的统一量化描述
[0014] By mapping keyword frequency vectors from transcribed call recordings to timestamp segments in customer repayment logs and aligning them with the timeline of external location tracking data, a fusion index sequence is generated. This computational method differs from simple timestamp alignment and overlay; instead, it performs vector mapping calculations between textual semantic keyness and repayment operation time windows. This allows the generated fusion index sequence to simultaneously reflect the temporal alignment between the characteristics of collection communication content and the customer's actual repayment behavior, transforming previously isolated multi-source data into quantifiable process coordination metrics. This fusion index sequence provides a numerical foundation with clear physical meaning for subsequent business stage identification, avoiding the information fragmentation and judgment delays caused by displaying multi-source data separately in traditional methods.
Smart Images

Figure CN122675136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing and visualization technology, specifically to a multi-source business data fusion analysis and visualization management system for debt collection operations. Background Technology
[0002] The debt collection process involves multi-source, heterogeneous business data, including call communication, repayment operations, and field visits. This data is interconnected in time, space, and semantics, reflecting collection progress and customer status from different perspectives. Existing debt collection management systems typically collect data from various business nodes at a fixed frequency, storing and analyzing the collected text, logs, and tracking data independently. Call recordings are transcribed only for compliance checks, repayment logs for accounting reconciliation, and field visit tracking for attendance management. There is a lack of effective integration mechanisms between these data sources, failing to create a unified quantitative description of the collaborative status of the collection process. As the business enters different stages, such as initial contact, negotiation, and repayment execution, the importance and rate of change of each data source differ significantly. The fixed-frequency collection method cannot dynamically adjust the data collection density according to the characteristics of each stage, resulting in insufficient granularity of information acquisition at critical stages and redundant information at stable stages, leading to a waste of transmission, storage, and computing resources. Meanwhile, existing data visualization methods are limited to the time-series display of a single data source, failing to integrate the results of multi-source data fusion with the division of business stages. This makes it difficult for managers to quickly perceive the evolution of business stage transitions and data collaboration relationships. How to construct a fusion indicator that reflects the degree of collaboration among multi-source business data on a timeline, and how to use this indicator to automatically divide business stages and drive dynamic adjustments to the collection frequency, are urgent problems to be solved in the multi-source data fusion analysis and visualization management of debt collection operations. Summary of the Invention
[0003] This invention provides a system for the fusion analysis and visualization management of multi-source business data in the debt collection process. By mapping the keyword frequency vector in the text sequence of call recordings to the timestamp segment of the customer repayment operation log, a fusion index sequence reflecting the degree of collaboration of multi-source data is generated. The system uses this index sequence to perform iterative time window division, automatically identify business stages and dynamically adjust the data collection frequency, thereby realizing in-depth fusion analysis of multi-source heterogeneous data and adaptive control of collection resources.
[0004] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a multi-source business data fusion analysis and visualization management system for debt collection operations. The system includes a data acquisition unit, a multi-source data fusion analysis unit, a multi-stage dynamic management unit, and a visualization unit. The data acquisition unit synchronously acquires raw business data sets at multiple business nodes based on a preset acquisition frequency. The raw business data sets include call recording transcription sequences, customer repayment operation logs, and external visit location trajectory data. The multi-source data fusion analysis unit performs triggered data operations based on business type identifiers on the raw business data sets to generate a fusion degree index sequence. The fusion degree index sequence is obtained by mapping the keyword frequency vectors in the call recording transcription sequences to the timestamp segments of the customer repayment operation logs. The multi-stage dynamic management unit generates business stage division results through iterative time window segmentation based on the fusion degree index sequence, and dynamically adjusts the acquisition frequency of the data acquisition unit based on the business stage division results. This increases data acquisition density during critical business stages to capture subtle changes, and reduces the acquisition frequency during stable stages to save system resources, achieving precise adaptation between the acquisition strategy and business dynamics. The visualization unit generates a visualization control map by mapping the business stage division results and the integration index sequence to a multi-dimensional coordinate axis, so that the integration status and stage evolution of the collection operation process are presented intuitively, making it easy for managers to quickly identify the current business situation.
[0005] As a preferred technical solution of the present invention, the multi-source data fusion analysis unit specifically extracts the sentence-level emotional polarity value and the frequency of occurrence of customer identity confirmation sentences from the call recording transcribed text sequence to generate a first business type identifier sequence; arranges the customer repayment operation log in chronological order of operation time, extracts the amount change range and operation interval duration for each operation to generate a second business type identifier sequence; extracts the dwell time and address type label of each location point in the external visit location trajectory data to generate a third business type identifier sequence; performs time axis overlap alignment processing on the first business type identifier sequence, the second business type identifier sequence, and the third business type identifier sequence, calculates the ratio of the number of overlapping identifiers to the total number of identifiers in each time window, and generates the fusion degree index sequence. Through the semantic-level fusion of the above multi-source heterogeneous data, the discrete business identifiers extracted from text, operation logs, and spatial trajectories are unified into the same time coordinate system for overlap measurement, so that the fusion degree index can truly reflect the degree of coordination of customers in the three dimensions of emotional response, action cooperation, and spatial contact during the collection process, overcoming the defects of single data source in representing collection progress as one-sided and lagging information.
[0006] Furthermore, the multi-stage dynamic control unit uses the moment when the index value in the fusion index sequence first exceeds a preset threshold as the first dividing point, and divides the fusion index sequence into an initial stage sub-sequence and subsequent stage sub-sequences; performs sliding window mean filtering on the subsequent stage sub-sequences to generate a smooth index curve; identifies the turning point on the smooth index curve where the rate of change of slope exceeds the fluctuation benchmark value, and uses the turning point as the second dividing point; combines the initial stage sub-sequence and the sub-sequences divided by the second dividing point to form the business stage division result. This processing method automatically discovers stage boundaries by utilizing the abrupt change characteristics and trend turning point features of the fusion index sequence, avoiding the subjectivity and lag of manual division, and enabling the stage division result to dynamically adapt to the actual business rhythm.
[0007] In this technical solution, the multi-stage dynamic management and control unit dynamically adjusts the collection frequency of the data collection unit based on the business stage division results. Specifically, this includes: extracting stage tags corresponding to each sub-sequence in the business stage division results, whereby the stage tags include the initial contact stage, the agreement negotiation stage, and the repayment execution stage; querying a preset collection frequency lookup table based on the stage tags to obtain the target collection frequency value corresponding to each stage tag; and sending the target collection frequency value to the data collection unit to replace the current collection frequency on each business node. Thus, the collection frequency is automatically increased during the agreement negotiation stage to finely monitor negotiation progress and potential repayment signals, while the frequency is appropriately reduced during the repayment execution stage to reduce redundant transmission, forming a differentiated collection mechanism that adapts to the collection operation rhythm.
[0008] As another preferred technical solution of the present invention, the visualization unit constructs a two-dimensional coordinate system with the time axis as the horizontal axis and the integration index sequence as the vertical axis; a line graph of the integration index sequence is plotted in the two-dimensional coordinate system; the time range of each sub-sequence in the business stage division result is extracted, and a background color block is superimposed on the corresponding time area of the line graph, the color of the background color block corresponding to the stage label in the business stage division result; the two-dimensional coordinate system and the background color block are combined to generate the visualization control map. This map integrates the dynamic trend of integration degree with the stage division background, allowing operators to instantly locate the current collection stage through the color block, and judge the rationality of stage switching by combining the changes in the integration degree curve, providing a clear decision-making basis for adjusting collection strategies.
[0009] Preferably, the system further includes a priority sorting unit and a resource scheduling unit. The priority sorting unit ranks each business node in descending order based on the index values in the integration index sequence, generating a node processing priority list. This list includes a unique identifier for each business node and its corresponding processing order number. The resource scheduling unit allocates collection resources to the corresponding business nodes according to the processing order number based on the node processing priority list. By converting the integration index into quantifiable processing priorities, cases with high integration and smooth collection progress receive priority resource allocation to accelerate conversion, while cases with low integration can be temporarily deferred or switched to other strategies, thereby achieving dynamic optimal allocation of limited collection resources.
[0010] Specifically, the priority sorting unit associates each business node with its most recent index value in the integration index sequence, generating a node index value mapping table; sorts the business nodes from largest to smallest according to the index values in the node index value mapping table, generating a sorted list of business nodes; assigns an incrementing positive integer starting from 1 to each business node in the sorted list of business nodes as the processing order number; combines the unique identifier of the business node with the processing order number to generate the node processing priority list. The resource scheduling unit obtains the total amount of collection operation resources and the minimum amount of resources required by each business node; accesses business nodes sequentially according to the processing order number in the node processing priority list from smallest to largest; allocates the minimum amount of resources to the currently accessed business node and updates the remaining total amount of collection operation resources; repeats the aforementioned allocation process until the remaining total amount is insufficient to allocate to the business node corresponding to the next processing order number. This resource scheduling mechanism ensures that limited collection manpower, outbound call channels, and other resources are always allocated to the node with the highest integration, dynamically reflecting the real-time needs of business progress and improving overall collection efficiency.
[0011] As a further improvement to the present invention, the system also includes an anomaly warning unit. The anomaly warning unit performs window periodic analysis on the fusion index sequence, extracts the length and magnitude of consecutively decreasing subsequences in the fusion index sequence, and generates an anomaly warning signal when the length of the consecutively decreasing subsequence exceeds a preset length threshold and the magnitude of the decrease exceeds a preset magnitude threshold. Thus, the system can not only present information at the current stage but also automatically detect abnormal patterns of continuously deteriorating fusion, promptly issuing warnings to collection managers, prompting attention and intervention for specific cases, and avoiding missing the optimal collection opportunity due to continuously declining customer cooperation.
[0012] Furthermore, the anomaly warning unit sets a sliding window length and sequentially extracts subsequences from the fusion index sequence; calculates the difference between adjacent index values in each subsequence, and marks subsequences with all negative differences as candidate decreasing subsequences; counts the number of elements in the candidate decreasing subsequences as the length of the continuous decreasing subsequence; and calculates the absolute value of the difference between the first and last elements in the candidate decreasing subsequences as the decreasing amplitude. This continuous decreasing identification method uses monotonically decreasing within the window as the discrimination condition, which can accurately capture the trend of deterioration of the fusion index rather than random fluctuations, reduce the probability of false alarms, and ensure the effectiveness and reliability of the warning.
[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0014] By mapping keyword frequency vectors from transcribed call recordings to timestamp segments in customer repayment logs and aligning them with the timeline of external location tracking data, a fusion index sequence is generated. This computational method differs from simple timestamp alignment and overlay; instead, it performs vector mapping calculations between textual semantic keyness and repayment operation time windows. This allows the generated fusion index sequence to simultaneously reflect the temporal alignment between the characteristics of collection communication content and the customer's actual repayment behavior, transforming previously isolated multi-source data into quantifiable process coordination metrics. This fusion index sequence provides a numerical foundation with clear physical meaning for subsequent business stage identification, avoiding the information fragmentation and judgment delays caused by displaying multi-source data separately in traditional methods.
[0015] An iterative time window segmentation is implemented using a series of integration indexes. The moment when the index value first exceeds a preset threshold is used as the first dividing point. Subsequent sequences are then subjected to sliding window mean filtering, and the turning point where the slope change rate exceeds the fluctuation benchmark value is identified as the second dividing point. This automatically segments the entire collection process into the initial contact stage, the agreement negotiation stage, and the repayment execution stage. Based on this business stage segmentation, the data collection frequency of the data acquisition unit is dynamically adjusted, applying differentiated collection densities to call recordings, repayment operation logs, and external visit trajectory data at different business stages. Higher collection frequencies are used during critical stages where index values rise rapidly or fluctuate dramatically, obtaining finer-grained process samples and improving the ability to capture negotiation evolution and repayment actions. Lower collection frequencies are used during stages where index values are smooth and stable, reducing the generation of invalid data and resource consumption. This mechanism of automatic closed-loop adjustment of collection frequency according to business stages ensures that system resource allocation and business information needs remain consistent across time, balancing data integrity and resource utilization efficiency. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0017] Figure 1 This is a schematic diagram of the structure of a multi-source business data fusion analysis and visualization control system for the debt collection process;
[0018] Figure 2 This is a flowchart of the multi-source data fusion analysis unit generating a sequence of fusion degree indicators;
[0019] Figure 3 This is a flowchart of business phase division and dynamic adjustment of collection frequency based on the integration degree index sequence;
[0020] Figure 4 This is a diagram showing the priority ranking of resource allocation for debt collection operations. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] See Figure 1 This invention provides a multi-source business data fusion analysis and visualization management system for debt collection operations, comprising a data acquisition unit, a multi-source data fusion analysis unit, a multi-stage dynamic management unit, and a visualization unit. The data acquisition unit synchronously acquires raw business data sets at multiple business nodes based on a preset acquisition frequency. These raw business data sets include call recording transcription sequences, customer repayment operation logs, and external visit location trajectory data. The multi-source data fusion analysis unit performs triggered data operations based on business type identifiers on the raw business data sets to generate a fusion degree index sequence. This fusion degree index sequence is obtained by mapping the keyword frequency vectors in the call recording transcription sequences to the timestamp segments of the customer repayment operation logs. The multi-stage dynamic management unit generates business stage division results based on the fusion degree index sequence through iterative time window segmentation, and dynamically adjusts the acquisition frequency of the data acquisition unit based on these results. The visualization unit generates a visualization management map by mapping the business stage division results and the fusion degree index sequence to multi-dimensional coordinate axes.
[0023] Example 1:
[0024] In specific implementation, please refer to Figure 2 The process by which the multi-source data fusion analysis unit performs triggered data operations based on business type identifiers on the original business data set to generate a fusion degree index sequence is achieved in the following way.
[0025] The multi-source data fusion analysis unit extracts sentence-level sentiment polarity values and the frequency of customer identification confirmation statements from the call recording transcription text sequence. For each sentence in the call recording transcription text sequence, the multi-source data fusion analysis unit calls a pre-trained text sentiment analysis model to output a sentence-level sentiment polarity value. A positive sentence-level sentiment polarity value indicates positive sentiment, and a negative value indicates negative sentiment. Simultaneously, the multi-source data fusion analysis unit matches each sentence against a preset list of keywords for customer identification confirmation statements, counting the frequency of these statements. The multi-source data fusion analysis unit combines the sentence-level sentiment polarity value and the frequency of customer identification confirmation statements for each sentence in chronological order according to the sentence timestamps to generate a first business type identifier sequence. Each element in the first business type identifier sequence contains a timestamp and a corresponding sentence-level sentiment polarity value and a frequency marker for the customer identification confirmation statement.
[0026] The multi-source data fusion analysis unit arranges customer repayment operation logs in chronological order and extracts the amount change range and operation interval for each operation. The customer repayment operation logs record the time, amount, and operation type for each repayment operation. For two consecutive operations, the multi-source data fusion analysis unit calculates the absolute value of the difference between the amount of the later operation and the amount of the earlier operation as the amount change range; it also calculates the time between the time of the later operation and the time of the earlier operation as the operation interval. For the first operation, the amount change range is set to the amount of the first operation itself, and the operation interval is set to zero. The multi-source data fusion analysis unit associates the amount change range and operation interval for each operation with the timestamp of that operation to generate a second business type identifier sequence. Each element in the second business type identifier sequence contains a timestamp and the corresponding amount change range and operation interval.
[0027] The multi-source data fusion analysis unit extracts the dwell time and address type tags for each location point in the external visit location trajectory data. The external visit location trajectory data includes location points reported by the location devices carried by the visitors at regular intervals. Each location point includes the reporting time, latitude and longitude coordinates, and address description information. The multi-source data fusion analysis unit calculates the dwell time for each location point based on the time interval between consecutive location points. The dwell time is the duration between the reporting time of the last location point and the reporting time of the current location point; if the current location point is the last location point in the trajectory, the dwell time is set to a preset default dwell time. The multi-source data fusion analysis unit converts the address description information into preset address type tags through an address resolution service. Address type tags include residential type, workplace type, and public area type. The multi-source data fusion analysis unit associates the dwell time and address type tags for each location point with the location point timestamp to generate a third business type identifier sequence. Each element in the third business type identifier sequence contains a timestamp and the corresponding dwell time and address type tag.
[0028] After generating the first, second, and third business type identifier sequences, the multi-source data fusion analysis unit performs time axis overlap alignment processing on these sequences. The time axis overlap alignment processing is performed as follows: the multi-source data fusion analysis unit sets a fixed-length time window, the length of which is pre-configured based on the density of business type identifiers; the multi-source data fusion analysis unit scans the time axis by sliding the time window, and within each time window, it counts the total number of occurrences of all business type identifiers in the first, second, and third business type identifier sequences, which is taken as the total number of identifiers. Simultaneously, the multi-source data fusion analysis unit counts the number of times at least two types of business type identifiers appear simultaneously in the first, second, and third business type identifier sequences within that time window, which is taken as the identifier overlap quantity. The identifier overlap quantity is calculated by taking the time point when a business type identifier appears in each sequence within the time window as an event point; when event points from different sequences exist at the same time or within a very small adjacent time difference, it is counted as one identifier overlap.
[0029] The multi-source data fusion analysis unit calculates the ratio of the number of overlapping identifiers to the total number of identifiers within each time window, generating a fusion degree index sequence. The fusion degree index sequence... The integration index value corresponding to each time window The calculation formula is:
[0030]
[0031] in, Indicating the first in the series of integration metrics The integration index value for each time window. Indicates the first The number of overlaps is marked within each time window. Indicates the first The total amount is identified within each time window. The time window number is incremented sequentially from 1.
[0032] The multi-source data fusion analysis unit arranges the fusion degree index values of each time window according to the time window sequence, forming a fusion degree index sequence. The multi-source data fusion analysis unit triggers the above data operation based on the business type identifier each time the original business data set is updated, updating the fusion degree index sequence.
[0033] Example 2:
[0034] In specific implementation, please refer to Figure 3 The multi-stage dynamic control unit generates business stage division results by iteratively dividing the data into time windows based on the integration degree index sequence.
[0035] The multi-stage dynamic control unit uses the moment when the value of the integration index in the sequence first exceeds a preset threshold as the first dividing point. The preset threshold is set based on the statistical characteristics of the integration index values at business process transition points in historical business data; the preset threshold is the median of the integration index values at the corresponding business state transitions in the historical integration index sequence. Starting from the first time window of the integration index sequence, the multi-stage dynamic control unit sequentially compares the integration index value of each time window with the preset threshold, determining the starting moment of the first time window where the integration index value is greater than or equal to the preset threshold as the first dividing point. The multi-stage dynamic control unit divides the subsequence of the integration index sequence from the sequence start point to the first dividing point into initial stage subsequences, and the portion after the first dividing point is designated as subsequent stage subsequences.
[0036] The multi-stage dynamic control unit performs sliding window mean filtering on subsequent stage sub-sequences to generate a smoothing index curve. The sliding window mean filtering is performed as follows: a sliding window length is set, which is obtained by multiplying the total number of elements in the subsequent stage sub-sequence by a preset proportional coefficient, with the preset proportional coefficient set to 0.1. When the result of multiplying the total number of elements in the subsequent stage sub-sequence by 0.1 is less than 3, the sliding window length is set to 3. The multi-stage dynamic control unit slides the sliding window along the subsequent stage sub-sequence, moving one element position at a time, and calculates the average of all fusion index values within the current sliding window as the smoothing index value. The smoothing index values calculated at each position in the subsequent stage sub-sequence are sequentially connected by covering each position with the sliding window to form the smoothing index curve.
[0037] After generating the smoothing index curve, the multi-stage dynamic control unit identifies the inflection point on the smoothing index curve where the rate of change of slope exceeds the fluctuation benchmark value, and uses this inflection point as the second boundary point. The rate of change of slope reflects the degree of change in the slope of the smoothing index curve, and the identification process is achieved by calculating the second difference between three adjacent points. For three consecutive points on the smoothing index curve, the coordinates are... , , ,in , , These are the smoothing index values for the corresponding time windows. , , These are the sequence number or time marker of the corresponding time window, and the calculation point of the multi-stage dynamic control unit. Rate of change of slope at point The calculation formula is:
[0038]
[0039] In the formula, Indicates the first on the smoothing index curve Rate of change of slope at each point location Indicates the first The smoothing index value corresponding to each time window Indicates the first The smoothing index value corresponding to each time window Indicates the first The smoothing index value corresponding to each time window Indicates the first The sequence number of each time window. Indicates the first The sequence number of each time window. Indicates the first The time window number. The volatility benchmark value is set based on the standard deviation of the absolute value of the rate of change of the slope at each point in the smoothing index curve. The volatility benchmark value is 1.5 times the standard deviation of the absolute value of the rate of change of the slope at all points in the smoothing index curve. When a certain point... Rate of change of slope at point When the absolute value exceeds the fluctuation benchmark value, the multi-stage dynamic control unit identifies the corresponding time position as a turning point. All identified turning points are then used as the second dividing point in sequence.
[0040] The multi-stage dynamic control unit combines the initial stage subsequence and the subsequences divided by the second boundary point into a business stage division result. The business stage division result consists of multiple consecutive subsequences, each subsequence corresponding to a time interval.
[0041] The multi-stage dynamic management unit dynamically adjusts the data collection frequency of the data acquisition unit based on the business stage segmentation results. The multi-stage dynamic management unit extracts the stage label corresponding to each sub-sequence in the business stage segmentation results. The correspondence of stage labels is determined by the sequential position of the business stage segmentation results: the stage label corresponding to the initial stage sub-sequence is the initial contact stage; the stage label corresponding to the first sub-sequence immediately following the initial stage sub-sequence is the agreement negotiation stage; and the stage labels corresponding to the remaining sub-sequences are the repayment execution stage. When the number of second boundary points exceeds one, the first second boundary point delineates the agreement negotiation stage, and all subsequent sub-sequences are marked as repayment execution stages.
[0042] The system queries a pre-defined collection frequency lookup table based on stage labels to obtain the target collection frequency value corresponding to each stage label. The pre-defined collection frequency lookup table stores the first collection frequency value for the initial contact stage, the second collection frequency value for the agreement negotiation stage, and the third collection frequency value for the repayment execution stage. The first, second, and third collection frequency values are pre-configured based on the different timeliness requirements of each stage, with the first collection frequency value being higher than the third, and the second collection frequency value falling between the first and third. The multi-stage dynamic management unit packages the time interval of each sub-sequence with the obtained target collection frequency value into a frequency adjustment instruction and sends it to the data acquisition unit. Upon receiving the frequency adjustment instruction, the data acquisition unit replaces the current collection frequency on each business node with the target collection frequency value within the corresponding time interval. When the time interval ends, the data acquisition unit automatically switches to the preset default collection frequency or receives a new frequency adjustment instruction.
[0043] Example 3:
[0044] In practice, the visualization unit generates a visual management and control map by mapping the business stage division results and integration index sequence to a multi-dimensional coordinate axis.
[0045] The visualization unit constructs a two-dimensional coordinate system with the time axis as the horizontal axis and the integration index sequence as the vertical axis. The time axis's scale covers the start time of the first time window in the integration index sequence to the end time of the last time window. The scale interval of the time axis is determined based on the number of time windows in the integration index sequence. When the number of time windows in the integration index sequence exceeds a preset dense display threshold, the scale interval of the time axis is automatically adjusted to twice the original scale interval. The vertical axis of the integration index sequence starts from zero, with an upper limit of 1.2 times the maximum integration index value in the integration index sequence. The scale interval of the vertical axis is set equally according to the numerical distribution range of the integration index sequence.
[0046] After constructing the two-dimensional coordinate system, the visualization unit plots a line graph of the integration index sequence within that system. The visualization unit uses the midpoint of each time window in the integration index sequence as the x-axis value and the corresponding integration index value for each time window as the y-axis value, determining the location of the data point for each time window in the two-dimensional coordinate system. The visualization unit connects the data points corresponding to adjacent time windows with straight line segments to form the line graph of the integration index sequence. For missing or outlier values in the integration index sequence, the visualization unit uses linear interpolation to supplement the corresponding y-axis values before plotting the line graph.
[0047] The visualization unit extracts the time range of each subsequence from the business phase segmentation result. The business phase segmentation result contains multiple subsequences, each subsequence corresponding to a phase label and a time interval, which is defined by the start and end times. The visualization unit iterates through each subsequence in the business phase segmentation result, obtaining the start time, end time, and phase label for each subsequence.
[0048] The visualization unit overlays background color blocks onto the corresponding time regions of the line chart. For each subsequence, the visualization unit uses the start and end times of the subsequence as the horizontal boundary of the background color block, and the lower and upper limits of the vertical axis of the two-dimensional coordinate system as the vertical boundary of the background color block, generating a rectangular area. The color of the background color block corresponds to the stage label in the business stage division results, and the color correspondence is determined by a preset stage label color mapping table. In the stage label color mapping table, the color value corresponding to the initial contact stage is semi-transparent blue, the color value corresponding to the agreement negotiation stage is semi-transparent yellow, and the color value corresponding to the repayment execution stage is semi-transparent green. The semi-transparency setting makes the background color block and the covered line chart lines visible simultaneously. The visualization unit draws the generated rectangular area on the lower layer of the line chart in a layer overlay manner.
[0049] The color assignment function used in the visualization unit construction stage label color mapping table to the stage labels and color values is:
[0050]
[0051] in, Indicates the first The color assignment results for each stage label This is the index number of the stage label in the stage label color map table. The value can be 1, 2 or 3, corresponding to the initial contact stage, the agreement negotiation stage and the repayment execution stage, respectively. Indicates the first The preset color vector corresponding to each stage label It is a four-dimensional vector composed of the red, green, and blue color channel components and the alpha channel component. For the initial contact phase, the preset color vector has a red channel component of 100, a green channel component of 150, a blue channel component of 255, and an opacity channel component of 100. During the protocol negotiation phase, the preset color vector has a red channel component of 255, a green channel component of 255, a blue channel component of 100, and an opacity channel component of 100. For the corresponding repayment execution phase, the preset color vector has a red channel component of 100, a green channel component of 255, a blue channel component of 100, and an opacity channel component of 100. The opacity channel component has a value range of 0 to 255, with a value of 100 representing approximately 60% opacity, ensuring that the lines in the underlying line chart are still legible.
[0052] The visualization unit combines a two-dimensional coordinate system and background color blocks to generate a visual control map. The combination process includes merging the completed line chart layer with all background color block layers, preserving the grid lines and scale labels in the two-dimensional coordinate system, adding a map title and axis labels, and generating the final visual control map. The output format of the visual control map is a vector graphics file or a bitmap file. Vector graphics file formats include scalable vector graphics formats, and bitmap file formats include portable network graphics formats.
[0053] Example 4:
[0054] In practice, the system also includes a priority sorting unit and a resource scheduling unit.
[0055] The priority sorting unit arranges each business node in descending order based on the index values in the integration index sequence, generating a node processing priority list. The node processing priority list contains a unique identifier for each business node and its corresponding processing order number.
[0056] The priority sorting unit associates each business node with its most recent metric value in the integration metric sequence, generating a node metric value mapping table. For each business node, the priority sorting unit obtains all integration metric values corresponding to that business node in the integration metric sequence, selecting the integration metric value with the latest timestamp as the most recent metric value. The priority sorting unit stores the unique identifier of each business node and its corresponding most recent metric value in key-value pairs, where the key is the unique identifier of the business node and the value is the most recent metric value. The set of all key-value pairs constitutes the node metric value mapping table.
[0057] After generating the node indicator value mapping table, the priority sorting unit sorts the business nodes from largest to smallest according to the indicator values in the table. The sorting process uses a comparison sorting method, comparing the most recent indicator values of each business node in the node indicator value mapping table pairwise, and rearranging the business nodes in descending order of their most recent indicator values to generate a sorted list of business nodes. When two or more business nodes have the same most recent indicator value, the priority sorting unit determines the order based on the lexicographical order of the unique identifiers of the business nodes.
[0058] The priority sorting unit assigns an incrementing positive integer, starting from 1, to each business node in the sorted list as its processing order number. The first business node in the sorted list is assigned processing order number 1, the second is assigned processing order number 2, and so on. The processing sequence number assigned to each business node is: .
[0059] The priority sorting unit combines the unique identifier of each business node with its processing sequence number to generate a node processing priority list. The combination method involves treating each business node's unique identifier and its corresponding processing sequence number as a tuple, and then arranging all tuples in ascending order of their processing sequence numbers to form the node processing priority list. Each tuple in the node processing priority list contains two fields: the first field is the processing sequence number, and the second field is the unique identifier of the business node.
[0060] After the node processing priority list is generated, the resource scheduling unit allocates the collection job resources to the corresponding business nodes in sequence according to the processing order number, based on the node processing priority list.
[0061] The resource scheduling unit obtains the total amount of collection operation resources and the minimum resource requirement for each business node. The total amount of collection operation resources is the total number of all collection operation resources currently available in the system. Collection operation resources include the sum of the number of field personnel and communication lines. The minimum resource requirement for each business node is the minimum resource configuration required to start collection operations at that business node. The minimum resource requirement is obtained by dividing the current number of pending cases at each business node by the preset maximum number of cases that can be processed per resource. When the calculated result is less than 1, the minimum resource requirement is set to 1.
[0062] The resource scheduling unit accesses the service nodes sequentially according to their processing order numbers in the node processing priority list, from smallest to largest. Starting with the service node with processing order number 1, the resource scheduling unit obtains the unique identifier of the currently accessed service node and queries the node processing priority list for the minimum resource quantity corresponding to that unique identifier.
[0063] The resource scheduling unit allocates a minimum amount of resources to the currently accessing business node. The allocation process involves the resource scheduling unit allocating resources equal to the minimum amount from the total collection task resources, marking these as allocated resources, and binding them to the unique identifier of the currently accessing business node. The resource scheduling unit then updates the remaining total amount of collection task resources by subtracting the minimum amount of resources already allocated to the currently accessing business node from the original total amount of collection task resources, resulting in the new remaining total amount.
[0064] The resource scheduling unit repeats the aforementioned allocation process, accessing the next business node in the node processing priority list one by one in ascending order of processing sequence number. It obtains the minimum resource amount corresponding to that business node, allocates that minimum resource amount from the remaining total amount of collection job resources, and updates the remaining total amount of collection job resources again. This allocation process is repeated until the remaining total amount of collection job resources is insufficient to allocate to the business node corresponding to the next processing sequence number; that is, the remaining total amount of collection job resources is less than the minimum resource amount required by the business node corresponding to the next processing sequence number. At this point, the resource scheduling unit stops allocating resources. All business nodes with processing sequence numbers after the currently allocated business nodes do not receive resource allocation.
[0065] The resource scheduling unit outputs the allocation results as a resource allocation scheme, which includes the unique identifier of the service node that has been allocated resources and the amount of resources obtained by each service node that has been allocated resources.
[0066] See Figure 4 In the graph, the horizontal axis represents the processing sequence number of the business nodes, increasing from 0 to 500. The vertical axis represents the integration index value of the corresponding business node, ranging from 0 to 1.0. The bars in the graph are sorted from largest to smallest integration index value, with the left bars being the tallest and having an integration index value close to 1.0, and the right bars gradually decreasing in height, with the lowest integration index value being approximately 0.2. The bar colors distinguish resource allocation status, with blue representing business nodes that have been allocated resources and gray representing business nodes that have not been allocated resources.
[0067] As described in Example 4, the priority sorting unit sorts the business nodes in the integration index sequence in descending order to generate a node processing priority list. The sorting in this figure reflects this process; nodes with higher integration indices on the left are sorted first and have smaller processing order numbers. The resource scheduling unit allocates collection operation resources to the business nodes sequentially according to the node processing priority list.
[0068] The blue bars in the diagram represent approximately 150 service nodes, indicating that the resource scheduling unit has allocated resources sequentially from 1 to approximately 150 according to their processing order numbers. The allocated resources meet the minimum resource requirements of each node. The gray bars starting to the right of the blue bars indicate that the resource scheduling unit has stopped allocating resources. The remaining resources are insufficient to meet the minimum resource requirements of service nodes with processing order numbers greater than approximately 150, therefore these service nodes have not received resource allocation.
[0069] Example 5:
[0070] In practical implementation, the system also includes an anomaly warning unit. The anomaly warning unit is used to perform window periodic analysis on the fusion index sequence, extract the length and magnitude of the continuous decreasing subsequence in the fusion index sequence, and generate an anomaly warning signal when the length of the continuous decreasing subsequence exceeds a preset length threshold and the magnitude of the decrease exceeds a preset magnitude threshold.
[0071] The anomaly warning unit sets a sliding window length and sequentially extracts subsequences from the fusion index sequence. The sliding window length is determined by multiplying the total number of elements in the fusion index sequence by a preset window length ratio coefficient. The preset window length ratio coefficient is set to 0.2, based on the fact that in the retrospective analysis of multiple historical anomalies, the minimum duration of a continuous downward trend consistently accounts for no less than 20% of the entire monitoring period. Therefore, 0.2 is chosen to ensure that the sliding window can cover the minimum span of the anomaly characteristics. When the result of multiplying the total number of elements in the fusion index sequence by 0.2 is less than 4, the sliding window length is set to 4. The anomaly warning unit extracts elements from the fusion index sequence using the sliding window length. The extraction method is as follows: starting from the first element of the fusion index sequence, each extraction extracts consecutive elements equal to the sliding window length as a subsequence. After each extraction, the starting position is moved one element to the right until the sum of the starting position and the sliding window length exceeds the total number of elements in the fusion index sequence, at which point the extraction process stops. All extracted element sets are arranged in the extraction order to form a subsequence set.
[0072] After obtaining the set of subsequences, the anomaly warning unit calculates the difference between adjacent index values in each subsequence and marks subsequences with all negative differences as candidate descent subsequences. For each subsequence in the set, starting from the first element, the anomaly warning unit calculates the difference between each subsequent element and the previous element, obtaining a set of differences. The anomaly warning unit performs a sign judgment on this set of differences; when the number of differences equals the number of elements in the subsequence minus one and all differences are negative, the subsequence is marked as a candidate descent subsequence.
[0073] After marking the candidate decreasing subsequences, the anomaly warning unit counts the number of elements in each candidate decreasing subsequence as the length of the consecutive decreasing subsequence. For each candidate decreasing subsequence, the anomaly warning unit directly obtains the total number of elements in the candidate decreasing subsequence as the length of the corresponding consecutive decreasing subsequence. When there are multiple candidate decreasing subsequences, the anomaly warning unit counts the length of the corresponding consecutive decreasing subsequence for each candidate decreasing subsequence separately.
[0074] The anomaly warning unit calculates the absolute value of the difference between the first and last elements of a candidate descending subsequence as the descending magnitude. For each candidate descending subsequence, the anomaly warning unit obtains the fusion index value of the first element in the subsequence, which is recorded as the starting value for calculating the descending magnitude; it also obtains the fusion index value of the last element in the subsequence, which is recorded as the ending value for calculating the descending magnitude. The anomaly warning unit calculates the absolute value of the difference between the starting and ending values for calculating the descending magnitude as the descending magnitude value corresponding to that candidate descending subsequence. The formula for calculating the descending magnitude value is as follows:
[0075]
[0076] in, Indicates the first The descent magnitude values corresponding to each candidate descent subsequence The index of the candidate descending subsequence is 1, and it is incremented sequentially. Indicates the first The fusion index value of the first element in each of the candidate descent subsequences. Mark the position of the first element in the candidate descending subsequence. Indicates the first The fusion index value of the last element in each candidate descending subsequence Mark the position of the last element in the candidate decreasing subsequence. This indicates the operation of taking the absolute value.
[0077] After calculating the length and magnitude of the consecutive decreasing subsequences for each candidate decreasing subsequence, the anomaly warning unit compares the length of the consecutive decreasing subsequences with a preset length threshold, and simultaneously compares the magnitude of the decrease with a preset magnitude threshold. The preset length threshold is set based on the length of the monitoring period and is half the length of the sliding window. The preset magnitude threshold is set based on the historical fluctuation range of the fusion index sequence and is set to 30% of the difference between the maximum and minimum values of the fusion index sequence in the most recent complete business cycle. When a candidate decreasing subsequence exists whose consecutive decreasing subsequence length exceeds the preset length threshold and whose magnitude of decrease exceeds the preset magnitude threshold, the anomaly warning unit generates an anomaly warning signal. The anomaly warning signal includes the start time of the abnormal decreasing trend, the length of the consecutive decreasing subsequences, the magnitude of the decrease, and the business node identifier to which the candidate decreasing subsequence belongs.
[0078] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A system for the fusion analysis and visualization management of multi-source business data in the debt collection process, characterized in that, The system includes: The data acquisition unit is used to synchronously acquire raw business data sets on multiple business nodes based on a preset acquisition frequency. The raw business data sets include call recording transcription text sequences, customer repayment operation logs, and external visit location trajectory data. The multi-source data fusion analysis unit is used to perform trigger-based data operations on the original business data set based on business type identifiers to generate a fusion degree index sequence. The fusion degree index sequence is obtained by mapping the keyword frequency vector in the call recording transcribed text sequence to the timestamp segment of the customer repayment operation log. A multi-stage dynamic management and control unit is used to generate business stage division results through iterative time window division processing based on the fusion index sequence, and dynamically adjust the collection frequency of the data collection unit based on the business stage division results. The visualization unit is used to generate a visual management and control map by mapping the business stage division results and the integration index sequence to a multi-dimensional coordinate axis.
2. The multi-source business data fusion analysis and visualization control system for debt collection operations according to claim 1, characterized in that, The multi-source data fusion analysis unit performs triggered data operations based on business type identifiers on the original business data set to generate a fusion degree index sequence, including: Extract the sentence-level sentiment polarity value and the frequency of occurrence of customer identity confirmation sentences from the call recording transcribed text sequence to generate a first business type identifier sequence; Arrange the customer repayment operation logs in chronological order of operation time, extract the amount change range and operation interval of each operation, and generate a second business type identifier sequence. Extract the dwell time and address type label of each location point from the external visit location trajectory data to generate a third business type identifier sequence; The first business type identifier sequence, the second business type identifier sequence, and the third business type identifier sequence are aligned by time axis overlap, and the ratio of the number of overlapping identifiers to the total number of identifiers in each time window is calculated to generate the fusion index sequence.
3. The multi-source business data fusion analysis and visualization control system for the collection process according to claim 2, characterized in that, The multi-stage dynamic management and control unit generates business stage division results based on the fusion degree index sequence through iterative time window division processing, including: The moment when the index value in the fusion index sequence first exceeds a preset threshold is taken as the first dividing point, and the fusion index sequence is divided into an initial stage subsequence and a subsequent stage subsequence. The subsequent stage subsequences are subjected to sliding window mean filtering to generate a smooth index curve; Identify the inflection point on the smoothed index curve where the rate of change of slope exceeds the fluctuation benchmark value, and use the inflection point as the second dividing point. The initial stage subsequence and each subsequence divided by the second dividing point are combined to form the business stage division result.
4. The multi-source business data fusion analysis and visualization control system for the collection process according to claim 3, characterized in that, The multi-stage dynamic management and control unit dynamically adjusts the data collection frequency of the data collection unit based on the business stage division results, including: Extract the stage label corresponding to each sub-sequence in the business stage segmentation result. The stage label includes the initial contact stage, the agreement negotiation stage, and the repayment execution stage. The target acquisition frequency value corresponding to each stage label is obtained by querying a preset acquisition frequency lookup table based on the stage label. The target acquisition frequency value is sent to the data acquisition unit to replace the current acquisition frequency on each service node.
5. The multi-source business data fusion analysis and visualization control system for the collection process according to claim 1, characterized in that, The visualization unit generates a visual management map by mapping the business stage division results and the integration index sequence to a multi-dimensional coordinate axis, including: Construct a two-dimensional coordinate system with the time axis as the horizontal axis and the fusion index sequence as the vertical axis; Plot a line graph of the fusion index sequence in the two-dimensional coordinate system; Extract the time range of each subsequence from the business stage division result, and overlay a background color block on the corresponding time area of the line chart. The color of the background color block corresponds to the stage label in the business stage division result. The two-dimensional coordinate system and the background color blocks are combined to generate the visual control map.
6. The multi-source business data fusion analysis and visualization control system for debt collection operations according to claim 1, characterized in that, The system also includes: The priority sorting unit is used to sort each business node in descending order according to the index value in the fusion index sequence and generate a node processing priority list. The node processing priority list includes the unique identifier of each business node and the corresponding processing order number. The resource scheduling unit is used to allocate collection operation resources to the corresponding business nodes in sequence according to the processing order number, based on the node processing priority list.
7. The multi-source business data fusion analysis and visualization control system for debt collection operations according to claim 6, characterized in that, The priority sorting unit sorts each business node in descending order according to the index values in the fusion index sequence, and generates a node processing priority list, including: Associate each business node with its most recent metric value in the fusion metric sequence to generate a node metric value mapping table; The business nodes are sorted from largest to smallest according to the index values in the node index value mapping table to generate a sorted list of business nodes. Assign an incrementing positive integer starting from 1 to each business node in the sorted business node list as the processing order number; The unique identifier of the business node is combined with the processing sequence number to generate the node processing priority list.
8. The multi-source business data fusion analysis and visualization control system for debt collection operations according to claim 7, characterized in that, The resource scheduling unit allocates collection task resources to the corresponding business nodes sequentially according to the processing order number, based on the node processing priority list, including: Obtain the total amount of resources required for the collection operation and the minimum amount of resources required for each business node; The business nodes are accessed sequentially from smallest to largest according to the processing order number in the node processing priority list. Allocate the minimum amount of resources to the currently accessed business node and update the remaining total amount of the collection job resources; Repeat the aforementioned allocation process until the remaining total amount is insufficient to allocate to the business node corresponding to the next processing sequence number.
9. The multi-source business data fusion analysis and visualization control system for debt collection operations according to claim 1, characterized in that, The system also includes: An anomaly warning unit is used to perform window periodic analysis on the fusion index sequence, extract the length and magnitude of the continuous decreasing subsequence in the fusion index sequence, and generate an anomaly warning signal when the length of the continuous decreasing subsequence exceeds a preset length threshold and the magnitude of the decrease exceeds a preset magnitude threshold.
10. The multi-source business data fusion analysis and visualization control system for the collection process according to claim 9, characterized in that, The anomaly warning unit performs window periodic analysis on the fusion index sequence, extracting the length and magnitude of continuously decreasing subsequences in the fusion index sequence, including: Set the sliding window length and sequentially extract subsequences from the fusion index sequence; Calculate the difference between adjacent index values in each subsequence, and mark the subsequences with all negative differences as candidate descent subsequences; The number of elements in the candidate decreasing subsequence is counted as the length of the continuous decreasing subsequence; The absolute value of the difference between the first and last elements in the candidate decreasing subsequence is calculated as the decreasing amplitude.