Intelligent communication assisting system for the elderly
By constructing multimodal feature data and a cone model of feature weight distribution, the problem of inaccurate evaluation in elderly call assistance systems in complex scenarios is solved, achieving more accurate call status perception and personalized call management.
Patent Information
- Application Number
- CN202511823223.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-12-05
AI Technical Summary
Existing call assistance systems for the elderly are unable to fully reflect the overall characteristics of the call content in complex call scenarios, and the feature weight allocation lacks dynamic correlation, resulting in inaccurate call status assessment and an inability to adapt to the personalized needs of different users or scenarios.
Multimodal feature data is constructed. The acquisition module acquires voice data streams in real time, the processing module performs risk keyword recognition and sentiment feature analysis, the optimization module calculates the feature weight distribution cone model, and the control module determines the call transmission control command based on the comprehensive evaluation results, thereby realizing the dynamic optimization of multi-dimensional fusion weight coefficients.
It improves the accuracy and adaptability of overall call assessment, provides more personalized call assistance decision-making, and enhances the effectiveness of call management and security protection.
Smart Images

Figure CN121281527B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to an intelligent communication assistance system for the elderly. Background Technology
[0002] With the widespread adoption of communication technology, the daily communication needs of the elderly are increasing. Due to age-related hearing loss, reduced comprehension, or mood swings, some elderly users may experience communication difficulties and incomplete information delivery during calls. To improve the call experience for the elderly, several assistive call systems have emerged, employing techniques such as voice enhancement, keyword recognition, or simple sentiment analysis to process call content.
[0003] However, some existing auxiliary systems rely on single-dimensional analysis methods when processing call content, such as identifying only specific keywords or simply classifying voice emotions. In complex call scenarios, such methods may not be able to fully reflect the overall characteristics of the call content, especially when the call content involves multiple semantic and emotional intertwines. Traditional methods often fail to fully combine multi-dimensional features, resulting in an inaccurate comprehensive assessment of the call status.
[0004] Furthermore, in terms of feature weight allocation and optimization, existing technologies mostly adopt fixed weights or allocation based on simple rules, lacking effective modeling of dynamic relationships between features. This approach may have insufficient adaptability when facing multimodal feature data, making it difficult to flexibly adapt to the personalized needs of different users or different call scenarios, thus affecting the effectiveness of the auxiliary system in practical applications. Summary of the Invention
[0005] The present invention aims to overcome the shortcomings of the prior art and provide an intelligent call assistance system for the elderly, which improves the system's perception, evaluation and control capabilities in complex call scenarios, and ensures the call experience and communication efficiency of elderly users.
[0006] To solve the above-mentioned technical problems, the basic concept of the technical solution adopted by the present invention is as follows:
[0007] A smart communication assistance system for the elderly, comprising:
[0008] The acquisition module is used to collect the user's voice data stream during the call in real time;
[0009] The processing module is used to parse the voice data stream to obtain parsed data; perform risk keyword recognition processing on the parsed data to obtain risk keyword recognition results; perform sentiment feature analysis processing on the parsed data to obtain sentiment feature analysis results; construct a three-dimensional feature distribution model based on the risk keyword recognition results and sentiment feature analysis results; calculate the volume parameters of the three-dimensional feature distribution model in a preset feature space; and combine the risk keyword recognition results, sentiment feature analysis results, and volume parameters to form multimodal feature data.
[0010] The optimization module is used to parameterize multimodal feature data to obtain an evaluation parameter set; construct a feature weight distribution cone model based on the evaluation parameter set; calculate the lateral area parameter of the feature weight distribution cone model, and fuse the lateral area parameter as a weight optimization parameter into the evaluation parameter set to obtain a fused parameter set; synthesize feature vectors from the fused parameter set to form multidimensional feature vectors; project the multidimensional feature vectors to obtain the mapping positions; perform spatial distribution analysis based on the mapping positions to obtain the corresponding multidimensional fusion weight coefficients; and calculate the comprehensive evaluation result based on the multidimensional fusion weight coefficients.
[0011] The control module is used to determine the call transmission control commands based on the comprehensive evaluation results.
[0012] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art.
[0013] This invention overcomes the limitations of traditional solutions that rely on single-dimensional analysis by constructing multimodal feature data that integrates risk keywords, emotional features, and three-dimensional distribution volume parameters. This multi-feature fusion mechanism can more comprehensively and three-dimensionally depict the overall state during a call, providing a more accurate and reliable data foundation for subsequent intelligent control.
[0014] This invention can transform the abstract weight allocation problem into a computable spatial geometric model. The method can dynamically parse and generate the optimal multidimensional fusion weight coefficients based on the specific characteristics of each call, effectively overcoming the poor adaptability of the fixed weight strategy, and making the final comprehensive evaluation result more in line with the actual situation of the current call scenario.
[0015] From three-dimensional feature distribution to weight optimization allocation, the spatial geometric modeling and calculation method adopted in this invention not only improves the representation ability of feature data, but also enables the entire system to have stronger contextual understanding and adaptive adjustment capabilities, thereby providing the elderly with more personalized call assistance decision-making that better meets their immediate needs. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. Some specific embodiments of this application will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings designate the same or similar parts or components. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0017] Figure 1 This is a schematic diagram of the intelligent call assistance system for the elderly according to the present invention.
[0018] Figure 2 This is a schematic diagram of the real-time acquisition process of user voice data during a call, as presented in this invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort should fall within the scope of protection of the present application.
[0020] The following embodiments of this application use an intelligent communication assistance system for the elderly as an example to illustrate the solution of this application in detail. However, this embodiment does not limit the scope of protection of this application.
[0021] like Figure 1 As shown, the present invention provides an intelligent communication assistance system for the elderly, comprising:
[0022] The acquisition module 11 is used to acquire the user's voice data stream during the call in real time;
[0023] Processing module 12 is used to parse the voice data stream to obtain parsed data; perform risk keyword recognition processing on the parsed data to obtain risk keyword recognition results; perform sentiment feature analysis processing on the parsed data to obtain sentiment feature analysis results; construct a three-dimensional feature distribution model based on the risk keyword recognition results and sentiment feature analysis results; calculate the volume parameters of the three-dimensional feature distribution model in a preset feature space; and combine the risk keyword recognition results, sentiment feature analysis results, and volume parameters to form multimodal feature data.
[0024] Optimization module 13 is used to parameterize multimodal feature data to obtain an evaluation parameter set; construct a feature weight distribution cone model based on the evaluation parameter set; calculate the lateral area parameter of the feature weight distribution cone model, and fuse the lateral area parameter as a weight optimization parameter into the evaluation parameter set to obtain a fused parameter set; synthesize feature vectors from the fused parameter set to form a multidimensional feature vector; project the multidimensional feature vector to obtain the mapping position; perform spatial distribution analysis based on the mapping position to obtain the corresponding multidimensional fusion weight coefficient; and calculate the comprehensive evaluation result based on the multidimensional fusion weight coefficient.
[0025] The control module 14 is used to determine the call transmission control command based on the comprehensive evaluation results.
[0026] In this embodiment of the invention, real-time acquisition of voice data streams ensures the timeliness and continuity of call voice data acquisition, providing reliable basic data for subsequent data processing and supporting the orderly conduct of multi-dimensional analysis. Through hierarchical parsing and feature extraction of the voice data stream, two core information categories—risk keywords and emotional features—are integrated. By constructing a three-dimensional feature distribution model and calculating volume parameters, the representational dimensions of feature data are expanded, strengthening the correlation between different types of features and improving the completeness and systematicity of multimodal feature data. Standardized integration of multimodal feature data is achieved through parameterized processing. Based on the construction of a feature weight distribution cone model and the calculation of lateral area parameters, the extraction and fusion of weight optimization parameters are realized. Through feature vector synthesis, projection, and spatial distribution analysis, the intrinsic correlation between parameters is deepened, making the determination of multi-dimensional fusion weight coefficients more consistent with data distribution patterns, thus improving the rationality and systematicity of parameter fusion. Based on the comprehensive analysis results from the previous stage, the determination of call transmission control commands becomes more scenario-adaptable, ensuring that intervention and control of calls meet actual needs and improving the effectiveness of call management and security protection.
[0027] In the intelligent call assistance system for the elderly described in this embodiment of the invention, the aforementioned acquisition module 11 acquires the user's voice data stream during the call in real time, including:
[0028] Step 1101: Receive the user's original voice signal through a microphone array; perform noise reduction processing on the original voice signal to obtain a noise-reduced voice signal; perform frame segmentation processing on the noise-reduced voice signal to divide the continuous voice signal into multiple voice frames.
[0029] Step 1102: Perform spatial location analysis on each speech frame, construct a three-dimensional spatial coordinate system based on the geometric layout of the microphone array, calculate the azimuth and elevation parameters of the sound source in the spherical coordinate system, and obtain a speech frame with spatial positioning information.
[0030] Step 1103: Perform pre-emphasis processing on the speech frame with spatial positioning information to obtain a pre-emphasis processed speech frame; perform windowing processing on the pre-emphasis processed speech frame to obtain a windowed speech frame.
[0031] Step 1104: Perform Fourier transform on the windowed speech frame to convert it into frequency domain feature data, and use the frequency domain feature data as the speech data stream.
[0032] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:
[0033] Step 1101 above first involves receiving the user's original voice signal via a microphone array. Specifically, multiple microphone units pre-deployed in the microphone array synchronously collect the user's voice signal during the call at a preset sampling frequency, thereby obtaining the original voice signal containing the user's voice information and environmental noise. After receiving the original voice signal, noise reduction processing is further performed on it. The specific process involves first collecting environmental noise samples from the current call environment, comparing these environmental noise samples with the original voice signal segment by segment, and identifying noise samples in the original voice signal that have the same characteristics as the environmental noise samples. The noise components are then filtered out from the original speech signal using signal filtering techniques to obtain a denoised speech signal. Based on the denoised speech signal, it is further processed by frame segmentation. Specifically, the preset frame length and frame shift parameters are determined according to the temporal characteristics of the speech signal. The frame length parameter is set to be long enough to completely retain a speech pitch period, and the frame shift parameter is set to be less than the frame length to ensure the signal continuity between adjacent frames. Then, according to the frame length and frame shift, the continuous denoised speech signal is divided into multiple independent speech frames along the time axis, and each speech frame retains the speech signal characteristics within the corresponding time period.
[0034] In step 1102 above, after obtaining multiple independent speech frames, spatial position analysis is performed for each speech frame. The specific calculation process begins by constructing a three-dimensional spatial coordinate system based on the actual geometric layout of the microphone array. First, the origin of this three-dimensional spatial coordinate system is determined; typically, the center of the microphone array is set as the origin. Then, the x-axis, y-axis, and z-axis are set, where the x-axis and y-axis form a horizontal plane parallel to the plane of the microphone array, and the z-axis is perpendicular to this horizontal plane and points upwards. Simultaneously, the specific coordinate values of each microphone unit in the array under this three-dimensional spatial coordinate system are recorded. After the three-dimensional spatial coordinate system is constructed, a spherical coordinate system mapping algorithm is used to calculate the azimuth and elevation parameters of the sound source. The process first calculates the time difference between each speech frame and the different microphone units in the microphone array. Based on this time difference and the preset sound wave propagation speed, the distance difference between the sound source and each microphone unit is calculated. Then, combined with the coordinate values of each microphone unit in the three-dimensional spatial coordinate system, a spatial geometric relationship between the sound source position and the position of each microphone unit is established. Based on this spatial geometric relationship, the rectangular coordinate parameters of the sound source in the three-dimensional spatial coordinate system are converted into spherical coordinate parameters, where the azimuth angle is the angle between the sound source in the horizontal plane and the positive x-axis, and the elevation angle is the angle between the sound source and the horizontal plane. Finally, the calculated azimuth and elevation angle parameters are associated with the corresponding speech frames to obtain speech frames with spatial positioning information.
[0035] In step 1103 above, after assigning spatial positioning information to the speech frame, pre-emphasis processing is performed on the speech frame with this spatial positioning information. First, the characteristic that high-frequency components of the speech signal are prone to attenuation during propagation is analyzed. For each speech frame with spatial positioning information, the amplitude value of the signal is adjusted sample by sample from the first signal sample to the last signal sample of the speech frame. The adjustment rule is to enhance the amplitude of signal components with frequencies higher than a preset threshold in the speech frame according to a preset amplitude gain coefficient, and keep the amplitude of signal components with frequencies lower than the preset threshold basically unchanged. Through this sample-by-sample processing, the pre-emphasized speech frame is obtained. After the pre-emphasis processing is completed, windowing processing is performed on the pre-emphasized speech frame. Specifically, a window function that meets the requirements of speech signal processing is selected. This window function must meet the requirement of smooth transition of the signal at the frame edge. Then, each amplitude value of the window function is multiplied point by point with the corresponding signal sample value of the pre-emphasized speech frame, so that the amplitude of the signal sample at the edge of the speech frame gradually attenuates, while the amplitude of the signal sample in the middle of the frame remains relatively stable. This reduces the spectral leakage problem caused by signal abrupt changes at the edge of the speech frame, and finally, the windowed speech frame is obtained.
[0036] In step 1104 above, after obtaining the windowed speech frame, a Fourier transform operation is performed on the windowed speech frame to convert it into frequency domain feature data. First, the number of Fourier transform points is determined according to the number of signal samples in the windowed speech frame to ensure that the number of transform points is not less than the number of samples in the windowed speech frame to guarantee transform accuracy. Then, all signal samples of the windowed speech frame are sequentially included in the transform process in chronological order. By performing frequency decomposition on the speech signal samples in the time domain, the frequency components contained in the speech frame are obtained. At the same time, the amplitude and phase values corresponding to each frequency component are calculated. Then, the amplitude and phase values of these frequency components are sorted out to select key parameters that can characterize the frequency features of the speech signal, such as the amplitude distribution of different frequency bands and peak frequency. These key parameters are integrated to form frequency domain feature data. Finally, the frequency domain feature data is used as the speech data stream required for subsequent processing.
[0037] In the intelligent call assistance system for the elderly described in this embodiment of the invention, the processing module 12 parses the voice data stream to obtain parsed data; performs risk keyword identification processing on the parsed data to obtain risk keyword identification results; performs emotional feature analysis processing on the parsed data to obtain emotional feature analysis results; constructs a three-dimensional feature distribution model based on the risk keyword identification results and the emotional feature analysis results; calculates the volume parameters of the three-dimensional feature distribution model in a preset feature space; and combines the risk keyword identification results, the emotional feature analysis results, and the volume parameters to constitute multimodal feature data, including:
[0038] Step 1201: Perform speech recognition processing on the speech data stream, convert the frequency domain feature data into a text sequence to obtain preliminary parsed data; perform semantic segmentation processing on the preliminary parsed data, and divide the text sequence into multiple semantic segments according to semantic coherence to obtain structured parsed data;
[0039] Step 1202: Perform word segmentation on the structured parsed data to obtain a keyword set; match and compare the keyword set with a preset risk keyword database to identify the successfully matched risk keywords;
[0040] Step 1203: Perform spatial mapping processing on the risk keywords, mapping each risk keyword to a preset two-dimensional semantic plane according to its semantic features, forming a spatial point set of risk keywords;
[0041] Step 1204: Based on the spatial point set of risk keywords, construct semantic association triangles between adjacent risk keywords, calculate the area value of each semantic association triangle, and determine the semantic association strength between risk keywords based on the area value.
[0042] Step 1205: Verify the validity of risk keywords based on semantic association strength and semantic paragraph context information to obtain valid risk keywords; classify and label the valid risk keywords, adjust the risk level labels according to semantic association strength, and obtain risk keyword identification results containing risk level labels.
[0043] Step 1206: Extract sentiment features from the structured parsed data. By analyzing the change trends of sentiment word density and sentiment intensity in semantic paragraphs, obtain sentiment feature analysis results including sentiment polarity value and sentiment intensity value.
[0044] Step 1207: Convert the risk level identifier in the risk keyword identification result into a risk intensity value, and use the sentiment polarity value and sentiment intensity value in the sentiment feature analysis result as the first dimension and the second dimension parameters, respectively.
[0045] Step 1208: Based on the risk intensity value, sentiment polarity value, and sentiment intensity value, map each semantic paragraph to a three-dimensional coordinate system to form a three-dimensional spatial point set;
[0046] Step 1209: Based on the distribution density of the three-dimensional spatial point set, construct the minimum convex polyhedron that encloses all spatial points, and use the minimum convex polyhedron as the three-dimensional feature distribution model.
[0047] Step 1210: Orthogonally project the minimum convex polyhedron along a direction perpendicular to the bottom surface of the preset feature space to obtain the projected contour; perform boundary extraction processing on the projected contour to identify the outer boundary curve of the projected contour.
[0048] Step 1211: Based on the outer boundary curve, calculate the minimum circumcircle of the projected profile, and use the radius of the minimum circumcircle as the radius parameter of the cylinder base; use the maximum height difference of the minimum convex polyhedron in the vertical direction as the cylinder height parameter.
[0049] Step 1212: Calculate the volume parameter of the three-dimensional feature distribution model in the preset feature space based on the radius parameter of the cylinder base and the height parameter of the cylinder, where the volume parameter is equal to the product of the area of the base circle and the height.
[0050] Step 1213: Synchronize and align the risk keyword identification results, sentiment feature analysis results, and volume parameters according to timestamps, and encapsulate them into a unified data structure to jointly constitute multimodal feature data.
[0051] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:
[0052] Step 1201 above first performs speech recognition processing on the speech data stream, specifically frequency domain feature data. During speech recognition, the system's built-in acoustic model is first invoked to perform acoustic feature matching on the frequency domain feature data, converting the frequency domain sound features into corresponding speech syllable sequences. Then, a preset language model is used to perform grammatical and semantic verification on the syllable sequences, correcting recognition errors. This converts the frequency domain feature data into a text sequence that can be directly semantically analyzed, thus obtaining preliminary parsed data. After obtaining the preliminary parsed data, semantic segmentation processing is further performed. Specifically, the text sequence in the preliminary parsed data undergoes grammatical structure analysis to identify punctuation marks indicating semantic pauses and conjunctions indicating topic transitions. Punctuation marks include periods and commas, while conjunctions include "but" and "in addition." Using these markers as demarcation points, the text sequence is divided into multiple semantically coherent paragraphs, each corresponding to a complete conversation semantic segment. Finally, structured parsed data is obtained, providing a well-organized data foundation for subsequent layered processing.
[0053] In step 1202 above, each semantic segment in the structured parsed data is first segmented into words. Specifically, a dictionary-based word segmentation algorithm is used, and the basic word segmentation dictionary is optimized by combining high-frequency words in the daily conversations of the elderly to ensure accurate recognition of common expressions used by the elderly. For example, colloquial expressions such as "transfer money" and "remit money" that the elderly may use can be effectively recognized. Then, the text in the semantic segment is broken down sentence by sentence, and the coherent sentences are split into independent word units. All word units are summarized to form a keyword set.
[0054] It should be noted that the preset risk keyword library used in the subsequent matching and comparison process is constructed as follows: First, based on publicly available cases related to fund security risks, typical publicly available cases of elderly people encountering fund or information security risks, and security risk warnings issued by authoritative security protection institutions, a comprehensive collection of risk-related words that frequently appear in call scenarios is compiled. Next, the collected words are categorized and organized according to the type of risk scenario, such as financial operation risk words, identity information risk words, and account security risk words. The financial operation category includes words related to fund transfer, and the identity information category includes words related to the leakage of personal identity information. The vocabulary, specifically for account security, covers terms related to changes in account permissions to facilitate matching risks in different scenarios. It also incorporates colloquial expressions commonly used by the elderly in daily conversations, supplementing the database with terms related to risky behaviors to avoid omissions due to differences in expression. For example, for risk terms like "safe account," it simultaneously includes similar expressions that the elderly might encounter, such as "protected account" or "dedicated account." Furthermore, this risk keyword database supports a regular update mechanism. The system can connect to the latest risk information released by official security platforms to promptly incorporate newly added risk terms into the database, adapting to constantly changing expressions related to risk behaviors and ensuring the timeliness and comprehensiveness of risk identification.
[0055] After obtaining the keyword set, the set is matched and compared with the aforementioned pre-constructed risk keyword library. The risk keyword library has been pre-stored with various high-frequency risk terms in scenarios such as fund security risks and personal information leakage risks, such as transfers in financial operations, ID numbers in identity information, and verification codes in account security. During the matching process, each word in the keyword set is compared with the word in the risk keyword library. Words that are completely identical or semantically equivalent are judged as successful matches, and all successfully matched words are recorded, thus identifying the successfully matched risk keywords.
[0056] In step 1203 above, spatial mapping is performed on each successfully matched risk keyword. In this process, a preset two-dimensional semantic plane needs to be constructed within the system first. The preset two-dimensional semantic plane needs to be combined with the characteristic patterns of risk keywords in the elderly call scenario and historical data support to ensure the accuracy and rationality of subsequent spatial mapping. Specifically, the plane construction first determines the definition of the core dimensions, and then refines the parameter settings of each dimension to form a complete plane system.
[0057] Regarding the dimensional definition and parameter preset of the two-dimensional semantic plane, the two core coordinate axes of the plane are first determined, namely the x-axis and the y-axis. The x-axis is preset as the category dimension of risk keywords. The category division of this dimension needs to be systematically sorted out based on common call risk scenarios among the elderly population. Common call risk scenarios include risks such as financial inducement, personal information leakage, and fraudulent service inducement. Specifically, risk keywords are divided into several major categories such as financial operation, identity information, account security, and fraudulent rights. Furthermore, a fixed and unique x-axis coordinate scale value is assigned to each category. For example, the financial operation category corresponds to x-axis coordinate value 1, the identity information category corresponds to x-axis coordinate value 2, the account security category corresponds to x-axis coordinate value 3, and the fraudulent rights category corresponds to x-axis coordinate value 4. This ensures that risk keywords of different categories have a clear positional distinction in the x-axis direction and avoids category confusion that leads to mapping deviation.
[0058] The y-axis is preset as the baseline dimension for the semantic relevance of keywords. The baseline value of this dimension is set based on the historical call data of the elderly accumulated by the system. Specifically, it is determined by analyzing the correlation data between risky keywords and risky events in historical calls. Risky events include verified risky calls and information security events. For example, the co-occurrence frequency of a certain risky keyword with other risky keywords in confirmed risky calls and the contribution of the keyword to the judgment of the risk event are statistically analyzed. Based on this data, the semantic relevance is divided into three baseline intervals: high, medium, and low. Each interval corresponds to a fixed y-axis coordinate value. For example, keywords with high co-occurrence frequency and large contribution correspond to a high baseline value on the y-axis, such as y-axis coordinate value 5; keywords with medium co-occurrence frequency and average contribution correspond to a medium baseline value on the y-axis, such as y-axis coordinate value 3; and keywords with low co-occurrence frequency and small contribution correspond to a low baseline value on the y-axis, such as y-axis coordinate value 1. This establishes the correspondence between semantic relevance and the y-axis coordinate, ensuring that the baseline value can accurately reflect the risk-related attributes of the keywords.
[0059] After completing the pre-setting of the above two-dimensional semantic plane, the semantic attributes of each risk keyword are analyzed. First, the risk category to which the keyword belongs is determined according to its specific meaning and actual application scenario, and then it is mapped to a pre-set fixed coordinate scale value on the x-axis. For example, the actual application scenario is that the word "transfer" is often used in the scenario of fund transfer. At the same time, according to the semantic relevance data of the keyword in historical call data, its corresponding y-axis benchmark value, that is, the semantic relevance benchmark value, is queried and determined. Then, based on the determined x-axis coordinates and y-axis coordinates, the specific coordinate points of each risk keyword are accurately located in the pre-set two-dimensional semantic plane. Finally, the coordinate points corresponding to all risk keywords are summarized to form a risk keyword spatial point set.
[0060] Step 1204 above involves constructing semantic association triangles between adjacent risk keywords based on the risk keyword spatial point set. First, the coordinate points in the spatial point set are sorted according to the order of occurrence of the risk keywords in the structured parsed data to determine the three adjacent coordinate points, i.e., the points corresponding to three consecutively occurring risk keywords. These three coordinate points are then connected sequentially to form a closed triangle, i.e., a semantic association triangle. Next, the area value of each semantic association triangle is calculated. Specifically, the side length of the triangle is calculated based on the coordinate difference by identifying the coordinates of the three vertices of the triangle, and then the area of the triangle is obtained using geometric calculation methods. After obtaining the area value, the semantic association strength between the risk keywords is determined based on the size of the area value. The system pre-defines the correspondence between area value and association strength; for example, the smaller the area value, the stronger the semantic association between the three keywords, and the larger the area value, the weaker the association, thus completing the quantitative determination of the semantic association strength.
[0061] Step 1205 above first verifies the validity of identified risk keywords by combining semantic association strength with semantic paragraph context information. Specifically, it first determines whether the semantic association strength of the risk keyword reaches the system's preset effective threshold. If it does not reach the threshold, it is initially determined to be invalid. For risk keywords that reach the threshold, it is further verified by combining the context content of the semantic paragraph in which they are located. For example, if the risk keyword is "transfer", and the context is that the child needs to pay tuition fees and needs to transfer money to the school account, then the reasonableness of the school account is considered to determine whether the keyword has actual risk significance. Keywords without actual risk are filtered out, and valid risk keywords are obtained. After obtaining valid risk keywords, they are classified and labeled. Category labels are added to each keyword according to risk categories, such as financial risk and information risk. At the same time, the risk level label is adjusted according to the semantic association strength. For example, keywords with high association strength correspond to high risk labels, and keywords with medium association strength correspond to medium risk labels. Finally, the risk keyword identification results containing risk level labels are obtained.
[0062] Step 1206 above involves extracting sentiment features from each semantic paragraph based on structured parsing data. First, a sentiment lexicon is pre-defined in the system, containing positive, negative, and neutral sentiment words. Positive sentiment words include "happy" and "reassured," while negative sentiment words include "anxious" and "afraid." Then, the number of positive and negative sentiment words in each semantic paragraph is counted, and the sentiment lexicon density is obtained by dividing the number of sentiment words by the total number of words in that paragraph. Simultaneously, the trend of sentiment intensity changes is analyzed. Specifically, each sentiment word is assigned a pre-defined sentiment intensity value according to its order of appearance in the semantic paragraph; for example, "anxious" corresponds to an intensity value of 3, and "afraid" corresponds to an intensity value of 4. These intensity values are then recorded sequentially, and the changes in intensity values as the paragraph progresses are observed, such as the increasing trend from calm to anxious. Combining the sentiment lexicon density and the sentiment intensity trend, the sentiment polarity (positive, negative, neutral) and corresponding sentiment polarity value of each semantic paragraph are determined, along with the specific numerical value of the sentiment intensity. Finally, the sentiment feature analysis results, including both the sentiment polarity value and the sentiment intensity value, are obtained.
[0063] Step 1207 above standardizes the feature parameters based on the risk keyword identification results containing risk level identifiers and the sentiment feature analysis results. First, the risk level identifiers in the risk keyword identification results are converted into risk intensity values. The system pre-sets the correspondence rules between levels and values, for example, low risk identifiers correspond to value 1, medium risk identifiers correspond to value 2, and high risk identifiers correspond to value 3, thus completing the vectorization of risk levels. At the same time, the sentiment polarity value and sentiment intensity value in the sentiment feature analysis results are determined as the first and second dimension parameters for subsequent modeling, respectively. Among them, the sentiment polarity value is assigned a specific value according to the preset rules, such as positive polarity corresponding to a positive value, negative polarity corresponding to a negative value, and neutral polarity corresponding to 0. The sentiment intensity value directly adopts the specific value obtained in step 1206 above, ensuring that different types of features are presented in a unified numerical form.
[0064] In step 1208 above, after parameter standardization, a three-dimensional coordinate system is constructed and spatial mapping of semantic paragraphs is achieved. First, three coordinate axes of the three-dimensional coordinate system are set: the x-axis corresponds to the risk intensity value, the y-axis corresponds to the first dimension parameter, i.e., the sentiment polarity value, and the z-axis corresponds to the second dimension parameter, i.e., the sentiment intensity value. At the same time, the numerical range of each coordinate axis is determined, i.e., a reasonable range is set based on historical data. Then, a spatial vector projection algorithm is used to treat the risk intensity value, sentiment polarity value, and sentiment intensity value corresponding to each semantic paragraph as three components of a three-dimensional vector. This vector is projected onto the preset three-dimensional coordinate system. According to the correspondence between the vector components and the coordinate axes, the specific coordinate points of the semantic paragraph in the three-dimensional coordinate system are determined. The coordinate points corresponding to all semantic paragraphs are summarized to form a three-dimensional spatial point set, realizing the spatial presentation of the multi-dimensional features of the semantic paragraph.
[0065] Step 1209 above is based on a three-dimensional spatial point set. The core is to construct a three-dimensional feature distribution model using the convex hull algorithm. This process requires phased completion of distribution density analysis, extreme point identification, and convex hull structure construction to ensure that the final model can completely represent the spatial distribution of three-dimensional features. The specific operations are as follows: First, a distribution density analysis of the three-dimensional spatial point set is performed. This analysis aims to provide a density basis for subsequent extreme point identification. For each feature point in the three-dimensional spatial point set, the straight-line distance in the three-dimensional coordinate system is calculated with each of the other feature points in the set. This is based on the x-axis risk intensity value, y-axis sentiment polarity value, and z-axis sentiment intensity value of the feature point to calculate the spatial straight-line distance between two points. Then, a distance threshold calibrated based on historical feature data is set. This threshold must effectively distinguish between dense and sparse regions of the point set. The number of other feature points contained within this distance threshold range for each feature point is counted. Based on the statistical results, the density distribution of the point set is divided, i.e., regions containing a large number of other feature points are dense regions, and regions containing a small number are sparse regions, thus determining the overall density pattern of the three-dimensional spatial point set.
[0066] After completing the distribution density analysis, the outermost extreme points of the point set are further identified. These extreme points are the core foundation for constructing the minimum convex polyhedron. The specific identification process is as follows: First, calculate the overall center coordinates of the three-dimensional point set, that is, take the average of the x-axis, y-axis, and z-axis coordinates of all feature points to obtain the x, y, and z values of the center coordinates; Second, calculate the spatial distance from each feature point to the center coordinates, and initially screen out a batch of candidate points that are far from the center coordinates; Third, combine the previous distribution density analysis results to perform a second screening of the candidate points, prioritizing the retention of those points. Candidate points located in sparse regions are more likely to be on the outer edge of the point set. Additionally, the maximum and minimum points in each of the three dimensions (x, y, and z) are further selected, such as the point with the highest risk intensity on the x-axis, the point with the most negative emotional polarity on the y-axis, and the point with the strongest emotional intensity on the z-axis, ensuring coverage of the boundary features of each dimension. In the fourth step, pairwise distances are calculated for the candidate points after the second selection, eliminating redundant points that are too close and all located in the same density region. Finally, all extreme points distributed on the outermost layer of the point set, the furthest from other points, and capable of completely defining the overall range of the point set are determined.
[0067] After identifying the extreme points, a convex hull construction operation can be performed based on these extreme points to gradually form a minimum convex polyhedron. The specific construction process is as follows: First, select four non-coplanar points from the identified extreme points as initial vertices. The criterion is that the fourth point is not on any plane formed by any three points, thus ensuring the spatial stability of the initial structure. Connect these four points pairwise to form a tetrahedral structure. This tetrahedron is the initial framework of the convex hull, and its four triangular faces are the initial outer surfaces of the convex hull. Second, traverse the three-dimensional space points. Focus on all feature points except the four vertices of the initial tetrahedron. For each feature point to be processed, perform an inside / outside judgment: calculate the spatial positional relationship between the point to be processed and each outer surface of the initial tetrahedron. Determine whether the point to be processed is inside or outside the surface by judging which side of the normal vector of the outer surface it is located on. If the point to be processed is inside any of the outer surfaces of the initial tetrahedron, it is determined that it is inside the convex hull, and no adjustment to the convex hull structure is required. If the point to be processed is outside at least one outer surface, it is determined that it is outside the convex hull, and the convex hull structure needs to be updated.
[0068] The third step is to perform a convex hull update operation on the points to be processed that are determined to be outside the convex hull: First, find all the outer surfaces of the initial tetrahedron that can see the point to be processed, i.e., the point to be processed is located outside these surfaces, and remove these surfaces from the outer surfaces of the convex hull; then, take the point to be processed as a new vertex and connect it to each edge of the removed surface to form a new triangular face, i.e., each edge corresponds to a new face, ensuring that the vertices of the new face are the point to be processed and the two endpoints of the edge; then, integrate these newly formed triangular faces with the outer surfaces of the initial tetrahedron that have not been removed, check and ensure that all faces are seamlessly connected and have no overlapping areas, forming a new closed polyhedron structure.
[0069] The fourth step is to repeat the above internal and external judgment and convex hull update operations until all feature points in the three-dimensional space point set have been traversed and processed. During this process, the volume of the polyhedron needs to be continuously verified. After each update, it is necessary to confirm whether the newly formed polyhedron can still completely enclose all feature points and has no extra faces or vertices. That is, all redundant structures that can be covered by other faces are eliminated. The final closed polyhedron is the smallest convex polyhedron that can completely enclose all space points and has the smallest volume.
[0070] Finally, the minimum convex polyhedron was determined as a three-dimensional feature distribution model. This model not only fully preserves the distribution range of the three-dimensional spatial point set in the three dimensions of risk intensity on the x-axis, sentiment polarity on the y-axis, and sentiment intensity on the z-axis, but also intuitively reflects the core distribution patterns such as the clustering area and dispersion of feature points corresponding to different semantic paragraphs in the three-dimensional space.
[0071] In step 1210 above, after the minimum convex polyhedron, orthogonal projection and boundary extraction operations are performed on the minimum convex polyhedron. The premise of orthogonal projection is based on a preset feature space. Therefore, before this, the preset construction of the preset feature space must be completed. This preset process must be closely connected with the construction logic of the three-dimensional feature distribution model in steps 1207 to 1209 above to ensure that the spatial attributes match the feature dimensions and provide a standardized spatial reference for subsequent projection operations.
[0072] Specifically, the pre-construction of the preset feature space needs to revolve around three core steps: First, determine the dimensional composition of the space. Combining the feature parameters determined in step 1207, namely risk intensity value, emotional polarity value, and emotional intensity value, the preset feature space is defined as a three-dimensional space. Its three coordinate axes correspond one-to-one with the x-axis, y-axis, and z-axis of the aforementioned three-dimensional coordinate system, that is, the x-axis represents the risk intensity value, the y-axis represents the emotional polarity value, and the z-axis represents the emotional intensity value. This ensures that the dimensions of the feature space are completely consistent with the dimensions of the multimodal features to be processed, avoiding feature mapping deviations due to dimensional misalignment. Second, determine the value range of each coordinate axis. Based on historical risk data and emotional performance data in elderly call scenarios, statistical analysis is performed. For example, the risk intensity value, combined with the risk level distribution in historical risk-related calls, is set to a range from 0 to 5, where 0 represents no risk. 5 represents extremely high risk. The emotional polarity value, combined with common emotional expressions, is set in the range of -3 to 3, where -3 represents extremely strong negative emotion and 3 represents extremely strong positive emotion. The emotional intensity value, combined with the intensity distribution of emotional words, is set in the range of 0 to 4, where 0 represents no obvious emotion and 4 represents extremely strong emotion, ensuring that the value range can fully cover all characteristic values that may appear in actual scenarios. Finally, the baseline attributes of the space are defined. The origin of the three-dimensional space is set as the coordinate point where the risk intensity value is 0, the emotional polarity value is 0, and the emotional intensity value is 0. This origin corresponds to the basic state of no risk, neutral emotion, and no obvious emotional intensity. At the same time, the positive and negative directions of each coordinate axis are defined, such as the positive direction of the x-axis representing increased risk intensity, the positive direction of the y-axis representing more positive emotional polarity, and the positive direction of the z-axis representing increased emotional intensity, thus forming a complete and standardized preset feature space.
[0073] Based on the aforementioned pre-defined feature space, the bottom surface of the feature space can be determined next. Combining the dimensional attributes of the feature space and the projection requirements, the bottom surface is usually set as the xy plane of the three-dimensional coordinate system. This plane simultaneously covers the two core basic feature dimensions of risk intensity and emotional polarity, which can preserve the spatial distribution information of these two key features to the greatest extent during the projection process, providing effective support for subsequent analysis of feature distribution patterns through projection contours. After determining the bottom surface, the minimum convex polyhedron is orthogonally projected along the direction perpendicular to the bottom surface, i.e., the z-axis direction, so that the minimum convex polyhedron in three-dimensional space forms a corresponding planar projection figure on the xy plane. This planar projection figure is the projection contour.
[0074] Subsequently, boundary extraction processing is performed on the projected outline. An edge detection method is used to traverse all pixels covered by the projected outline. By comparing the grayscale value of each pixel with the grayscale value of the surrounding background area pixels, pixels with significant differences in grayscale value are identified. These pixels are the boundary points between the projected outline and the background area, that is, the pixels at the edge of the outline. Then, according to the spatial order of these edge pixels in the projected outline, they are connected in sequence to form a continuous and closed curve. This curve is the outer boundary curve of the projected outline. The specific range boundary of the projected outline can be clearly defined by this outer boundary curve.
[0075] Step 1211 above determines the key parameters required for cylinder volume calculation based on the identified outer boundary curve of the projected profile and the minimum convex polyhedron. First, the minimum circumcircle of the projected profile is calculated. Specifically, by traversing all points on the outer boundary curve, the circle with the smallest radius that completely encloses the curve is found, and its radius is measured and used as the cylinder's base radius parameter. Simultaneously, the maximum height difference of the minimum convex polyhedron in the vertical direction (z-axis direction) is calculated. Specifically, the coordinate values of all vertices of the minimum convex polyhedron on the z-axis are identified, the maximum and minimum coordinate values are found, and the difference between them is calculated. This difference is the cylinder height parameter, ensuring that subsequent volume calculations accurately reflect the actual spatial scale of the 3D feature distribution model.
[0076] In step 1212 above, based on the cylinder base radius parameter and cylinder height parameter, the volume parameter of the three-dimensional feature distribution model in the preset feature space is calculated. First, according to the cylinder base radius parameter, the area of the cylinder base circle is obtained using the circle area calculation formula. This area can reflect the size of the projection range of the three-dimensional feature distribution model on the xy plane. Then, the calculated base circle area is multiplied by the cylinder height parameter, and the product is the volume parameter occupied by the three-dimensional feature distribution model in the preset feature space. This parameter can quantitatively characterize the spatial scale of the three-dimensional feature distribution model and provide a quantitative indicator of spatial dimension for subsequent multimodal feature data.
[0077] Steps 1213 above first collect the risk keyword identification results, sentiment feature analysis results, and volume parameters, and simultaneously extract the timestamp corresponding to each data point. This timestamp is consistent with the generation time of the semantic segment during the call and is obtained from the system call log. Then, the three types of data are synchronized and aligned according to the timestamps. The risk keyword identification results, sentiment feature analysis results, and volume parameters corresponding to the same timestamp are associated and matched to ensure that the three types of data are consistent in the time dimension and to avoid data association deviations caused by time misalignment. Finally, the aligned three types of data are encapsulated into a unified data structure, such as a system-defined structure or a standardized data frame format. This data structure includes risk information fields, sentiment feature fields, and volume parameter fields. The three types of data together constitute multimodal feature data.
[0078] In the intelligent call assistance system for the elderly described in this embodiment of the invention, the optimization module 13 performs parameterization processing on the multimodal feature data to obtain an evaluation parameter set; constructs a feature weight distribution cone model based on the evaluation parameter set; calculates the lateral area parameter of the feature weight distribution cone model, and integrates the lateral area parameter as a weight optimization parameter into the evaluation parameter set to obtain a fused parameter set; synthesizes feature vectors from the fused parameter set to form a multidimensional feature vector; projects the multidimensional feature vector to obtain a mapping position; performs spatial distribution analysis based on the mapping position to obtain the corresponding multidimensional fusion weight coefficient; and calculates the comprehensive evaluation result based on the multidimensional fusion weight coefficient, including:
[0079] Step 1301: Perform feature normalization processing on the multimodal feature data, convert the risk level identifier in the risk keyword identification result, the sentiment polarity value and sentiment intensity value in the sentiment feature analysis result, and the volume parameter into standardized values respectively, to obtain a set of normalized feature values;
[0080] Step 1302: Based on the normalized feature value set, extract the main feature components and use the main feature components as the initial parameters of the evaluation parameter set;
[0081] Step 1303: Based on the initial parameters of the evaluation parameter set, construct a three-dimensional parameter distribution space and map each evaluation parameter set to the corresponding coordinate point in the three-dimensional parameter distribution space;
[0082] Step 1304: Based on the distribution characteristics of all coordinate points in the three-dimensional parameter distribution space, identify the center point and boundary range of the evaluation parameter set distribution;
[0083] Step 1305: Construct a cone model of feature weight distribution by taking the center point of the evaluation parameter set distribution as the cone vertex and the boundary range as the cone base contour.
[0084] Step 1306: Extract geometric parameters from the feature weight distribution cone model to obtain the cone height parameter and base radius parameter;
[0085] Step 1307: Calculate the lateral area parameter of the feature weight distribution cone model based on the cone height parameter and the base radius parameter, where the lateral area parameter reflects the coverage and concentration of the parameter distribution;
[0086] Step 1308: Based on the lateral area parameter, determine the dynamic adjustment coefficient of each evaluation parameter, wherein the dynamic adjustment coefficient is inversely proportional to the lateral area parameter;
[0087] Step 1309: Multiply the dynamic adjustment coefficient with the corresponding parameter in the evaluation parameter set to obtain the adjusted parameter value;
[0088] Step 1310: The adjusted parameter values are fused with the lateral area parameters to obtain a fused parameter set.
[0089] Step 1311: Classify and arrange the parameters in the fusion parameter set according to the risk dimension, emotional dimension and spatial dimension to form a multidimensional feature vector with a three-dimensional structure;
[0090] Step 1312: Perform spherical coordinate transformation on the multidimensional feature vector to convert the multidimensional feature vector in the Cartesian coordinate system into the spherical coordinate transformation result. The spherical coordinate transformation result includes radial distance, azimuth angle and elevation angle parameters in the spherical coordinate system.
[0091] Step 1313: Based on the spherical coordinate transformation result, project the multidimensional feature vector onto the preset unit sphere to obtain the spherical projection point;
[0092] Step 1314: Convert the spherical coordinates of the spherical projection point back to Cartesian coordinates to obtain the mapped position coordinates on the unit sphere, which is used as the mapped position;
[0093] Step 1315: Perform spatial relationship analysis between the mapped position and the preset safe area spherical model, and calculate the spherical distance from the mapped position to the boundary of the safe area;
[0094] Step 1316: Based on the spherical distance, calculate the coverage area ratio of the region where the mapping location is located, and use the coverage area ratio as the risk weight coefficient;
[0095] Step 1317: Based on the distribution density of the mapped positions on the sphere, identify high-risk and low-risk areas to obtain the sentiment weight coefficients;
[0096] Step 1318: Calculate the comprehensive assessment result based on the risk weight coefficient and the sentiment weight coefficient, where the comprehensive assessment result reflects the call risk level.
[0097] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:
[0098] In step 1301 above, firstly, standardization rules are determined for each type of parameter in the multimodal feature data. For risk level identifiers in the risk keyword identification results, such as low risk, medium risk, and high risk, the textual level identifiers are first converted into initial values according to the system's preset level mapping rules. For example, low risk corresponds to 1, medium risk corresponds to 2, and high risk corresponds to 3. Then, the initial values are converted to a standardized range of 0 to 1 using a linear mapping method. For sentiment polarity values in the sentiment feature analysis results, such as the range of -3 to 3 (negative values represent negative sentiment, positive values represent positive sentiment, and 0 represents neutral), the Min-Max standardization algorithm is directly used, based on historical data. Using the extreme values of polarity (e.g., minimum -3, maximum 3) in the statistical analysis of emotional data from annual phone calls as a benchmark, the actual polarity values are scaled to the range of 0 to 1. For emotional intensity values, such as the range of 0 to 4, the Min-Max standardization algorithm is also used to scale the range of 0 to 1 based on the historical extreme values (0 and 4). For volume parameters, the maximum and minimum values of the volume parameters in the historical three-dimensional feature distribution model are first statistically analyzed, and then the current volume parameters are mapped to the range of 0 to 1 through linear transformation. After completing the standardization transformation of various parameters, all standardized values are classified and organized according to parameter type to form a set of normalized feature values that includes risk, emotional, and spatial standardized parameters.
[0099] In step 1302 above, firstly, principal component analysis is used to perform variance analysis on all parameters in the normalized eigenvalue set to calculate the variance contribution of each parameter, i.e., the explanatory power of the parameter for the overall feature differences. Then, a variance contribution threshold calibrated based on historical evaluation data is set, usually set to a cumulative variance contribution of 85% to ensure that core feature information is retained. Parameters are selected in descending order of variance contribution. Parameters with variance contribution higher than the individual parameter threshold (e.g., 10%) are retained first, and then parameters that can make the cumulative variance contribution reach 85% are selected from the remaining parameters and used as the main feature components. Finally, these main feature components are arranged in the order of risk-related parameters, sentiment-related parameters, and spatial-related parameters to form the initial parameters of the evaluation parameter set, ensuring that the initial parameters cover the core features and have a clear categorical structure.
[0100] In step 1303 above, firstly, the coordinate axes of the three-dimensional parameter distribution space are defined. The x-axis corresponds to the main risk-related feature components in the evaluation parameter set, such as the standardized risk level parameter; the y-axis corresponds to the main emotion-related feature components, such as the standardized emotion intensity parameter; and the z-axis corresponds to the main spatial-related feature components, such as the standardized volume parameter. Simultaneously, a value range of 0 to 1 is set for each coordinate axis, consistent with the range of normalized feature values. Then, for each evaluation parameter set, i.e., each group of main feature components corresponding to a certain semantic segment, its risk-related parameter values are extracted as x-axis coordinates, emotion-related parameter values as y-axis coordinates, and spatial-related parameter values as z-axis coordinates. These three coordinate values are combined to determine a unique corresponding coordinate point in the three-dimensional parameter distribution space, completing the mapping of all evaluation parameter sets to spatial coordinate points and forming an intuitive parameter space distribution pattern.
[0101] Step 1304 above involves conducting distribution characteristic analysis on three-dimensional spatial coordinate points to identify the center point and boundary range. The specific process is as follows: First, calculate the center point of the evaluation parameter set distribution. Take the arithmetic mean of the x-axis, y-axis, and z-axis coordinate values of all coordinate points to obtain the x-coordinate, y-coordinate, and z-coordinate of the center point. This center point reflects the central tendency of the parameter distribution. Then, determine the boundary range by processing the three dimensions of x-axis, y-axis, and z-axis separately. In the x-axis dimension, select the maximum and minimum values among all coordinate point x-values to form the boundary interval in the x-axis direction. Similarly, select the maximum and minimum values in the y-axis direction and the z-axis direction to form the boundary intervals in the y-axis and z-axis directions, respectively. During this process, outliers in the coordinate points need to be removed. Outliers are judged based on values that deviate from the average value of the dimension by more than three standard deviations to avoid interference from outliers on the boundary range. Finally, the boundary intervals of the x-axis, y-axis, and z-axis together constitute the three-dimensional boundary range of the evaluation parameter set, which can completely encompass all non-outlier coordinate points.
[0102] In step 1305 above, based on the identified center point and boundary range, constructing the feature weight distribution cone model requires first training the model using historical data to ensure the adaptability of geometric parameters and weight representations. Then, it is implemented using a standardized geometric modeling process. The specific operation needs to be carried out in three stages: model training, geometric construction, and system implementation, ensuring that the model accurately reflects the parameter distribution characteristics and stably supports subsequent weight optimization. The detailed process is as follows:
[0103] First, the feature weight distribution cone model is trained. This training aims to determine a reasonable benchmark for geometric parameters, such as vertex offset threshold, base contour scaling ratio, and height calculation coefficient. This avoids the model failing to correlate parameter distribution and weight features due to arbitrary setting of geometric parameters. The specific training process is as follows: Step 1: Prepare training data by selecting multimodal feature data of elderly people's calls accumulated in the past 6 months, covering call samples of low, medium, and high risk levels. The sample size should be no less than 5000 to ensure coverage of various call scenarios. These samples are then processed one by one according to steps 1301 to 1304 to obtain the center point coordinates, three-dimensional boundary range, and other parameters corresponding to each sample. Based on manually labeled actual risk assessment results, these data are integrated into a training dataset; simultaneously, 30% of the training dataset is allocated as a validation set for model validation. The second step involves generating an initial training model. For each sample in the training dataset, a temporary cone model is constructed according to the initial rules: the sample's center point is the vertex, the projection of the sample's 3D boundary range onto the xy-plane is the base contour, and the z-axis direction is the height direction. The geometric parameters of each temporary model, such as vertex coordinates, the length and width of the base rectangle, and the height, are recorded. The third step involves calculating the model error. The geometric parameters of each temporary model are correlated with the actual risk assessment results of the corresponding sample, and the error index is defined as the degree of concentration of parameter distribution represented by the model's geometric parameters. The matching degree with the actual risk assessment results is considered. For example, if the temporary model's base contour is too large for a high-risk sample, the representation parameters are scattered, but the actual risk assessment requires concentrated parameters to highlight risk characteristics, resulting in a high judgment error. The average error of the initial model is obtained by statistically analyzing the error values of all training samples. The fourth step involves iterative optimization. For samples with high errors, the geometric parameters of the temporary model are adjusted: if the base contour is too large, causing error, the base contour is reduced according to the correspondence between the risk level and the base scaling ratio in the training set (e.g., a scaling ratio of 0.8 for high-risk samples, 0.9 for medium-risk samples, and 1.0 for low-risk samples); if the height value is unreasonable, the height calculation coefficient is adjusted based on the standard deviation of the sample's z-axis boundary range. If the standard deviation is large, the height coefficient is set to 1.2; if the standard deviation is small, it is set to 0.9. If the vertex is offset, the vertex coordinates are finely adjusted in the offset direction corresponding to the actual risk characteristics of the sample. For example, the vertex of a high-risk sample is offset by 5% in the high-risk direction of the x-axis. In the fifth step, the model parameters are validated and fixed. The iteratively adjusted model parameters are applied to the validation set, and the average error of the validation set is calculated. If the error is lower than the preset acceptable threshold, such as the average error being less than or equal to 5%, the current geometric parameter benchmark is fixed as the system default parameter. The current geometric parameter benchmark includes the vertex offset threshold, the bottom scaling ratio, and the height calculation coefficient. If the error does not meet the standard, the above iterative optimization steps are repeated until the validation set error meets the requirements, and the model training is completed.
[0104] After training, based on the center point and boundary range identified in step 1304, and combined with the trained and fixed geometric parameter benchmarks, the core components of the cone model are determined: First, the cone vertex is determined. Based on the center point obtained in step 1304, if the risk tendency of the current call sample is high risk (where the risk tendency of the current call sample is based on the risk keyword identification result in step 1205), the vertex coordinates are finely adjusted towards the high-risk direction on the x-axis according to the trained and fixed vertex offset threshold. If it is low risk, no offset is needed, and the finally determined point is taken as the cone vertex. This vertex reflects both the core position of the parameter distribution in step 1304 and the associated risk weight features are adjusted through the offset after training. Second, the cone base contour is determined. First, the x-axis boundary interval and y-axis boundary interval of the three-dimensional boundary range obtained in step 1304 are extracted, and the x-axis interval length and y-axis interval length are calculated. Then, based on the risk level of the current call sample (i.e., the initial risk level in step 1205), the vertex coordinates are adjusted towards the high-risk direction on the x-axis. The first step involves scaling the x-axis and y-axis intervals using the trained scaling factor. For example, high-risk samples are scaled by 0.8 times, and medium-risk samples by 0.9 times. Then, on the xy-plane (risk and sentiment dimension plane), rectangles are drawn using the scaled x-axis and y-axis intervals. These rectangles form the base outline of the cone, ensuring the base size matches the risk weight requirements. Next, the cone's height direction and value are determined. The height direction is set to the z-axis direction perpendicular to the xy-plane, conforming to the trained height direction benchmark. The z-axis coordinate of the plane containing the base is set to the minimum z-axis value of the three-dimensional boundary range in step 1304. The z-axis coordinate of the vertex is based on the average z-axis value of the three-dimensional boundary range in step 1304, adjusted in conjunction with the trained height calculation coefficient. If the z-axis standard deviation is large, it is multiplied by 1.2; if the standard deviation is small, it is multiplied by 0.9. The difference between the two is the cone height, which reflects the concentration of parameters in the spatial dimension and their weight correlation.
[0105] Finally, the pyramid model is constructed and stored according to the system's preset geometric modeling rules, completing the model implementation process: Retrieving relevant data from the trained and fixed geometric modeling parameter library, the system uses the previously determined vertex coordinates, the coordinates of the four corner points of the base outline rectangle, and the height value as the basis for construction; the system executes geometric construction in a preset order, first connecting the four corner points of the base outline sequentially to form a closed rectangular base, then connecting the vertices to the four corner points of the base one by one to form four lateral triangles, ultimately forming a closed quadrangular pyramid structure; simultaneously, all geometric parameters of the quadrangular pyramid are automatically recorded, including vertex coordinates, coordinates of each corner point of the base, height value, and base side length, and these parameters are stored in a preset format in the model parameter file for easy retrieval of geometric parameters later; thus, the feature weight distribution pyramid model is constructed.
[0106] In step 1306 above, the cone height parameter is first calculated. This height is the vertical distance from the cone's vertex to the plane containing the base. The z-axis coordinate of the plane containing the base is set to the minimum z-axis value of the three-dimensional boundary range, and the z-axis coordinate of the vertex is the average z-axis value of the three-dimensional boundary range. The difference between the two is the cone height. Next, the base radius parameter is calculated. Since the base outline is a rectangle on the xy plane, the radius of the rectangle's circumcircle must be calculated first. The center of the rectangle is determined, which is the intersection of the rectangle's diagonals. The x-coordinate of this center is the midpoint of the x-axis boundary range, and the y-coordinate is the midpoint of the y-axis boundary range. Then, the length of the rectangle's diagonal is calculated, which is the distance from one corner point to the opposite corner point. Half the length of the diagonal is the radius of the rectangle's circumcircle, and this radius is the cone's base radius parameter. During the extraction process, it is necessary to ensure that the units of all parameters are consistent with the coordinate units of the three-dimensional parameter distribution space, providing a unified scale basis for subsequent lateral area calculations.
[0107] Step 1307 above first determines the calculation logic of the lateral surface area. Since the cone is a square pyramid, its lateral surface area is the sum of the areas of the four lateral triangles. The area of each lateral triangle needs to be calculated by first calculating the height of the triangle, which is the projection of the generatrix of the cone onto the lateral surface. First, calculate the generatrix of the cone, which is the spatial distance from the vertex to a corner of the base rectangle. The generatrix length is determined based on the geometric relationship between the base radius (radius of the circumcircle of the rectangle) and the height of the cone. Then, for each lateral triangle, take one side of the base rectangle as the base and the projection of the generatrix length onto that side as the height, which is the height of the lateral triangle, and calculate the area of a single lateral triangle. Add the areas of the four lateral triangles to obtain the lateral surface area parameter of the cone model. It should be noted that this lateral surface area parameter is directly related to the coverage and concentration of the parameter distribution. The larger the base radius (wider coverage) or the longer the generatrix length (more dispersed parameter distribution), the larger the lateral surface area parameter; conversely, the smaller the lateral surface area parameter, thus achieving a quantitative representation of the parameter distribution characteristics through the lateral surface area.
[0108] Step 1308 above, based on the lateral area parameter, determines the dynamic adjustment coefficients for each assessment parameter. The specific process is as follows: First, a baseline value for the lateral area is set. This baseline value is derived from statistical data on the parameter distribution of historical elderly people's calls, specifically the average value of the lateral area parameter in the historical data, representing the normal state of parameter distribution. Then, an inverse relationship is established between the dynamic adjustment coefficient and the lateral area parameter. The formula for calculating the dynamic adjustment coefficient is set as: Dynamic adjustment coefficient = Baseline value of lateral area / Current lateral area parameter. This ensures that the larger the current lateral area parameter, i.e., the more dispersed the parameter distribution, the smaller the dynamic adjustment coefficient; the smaller the current lateral area parameter... The smaller the value, the more concentrated the parameter distribution, and the larger the dynamic adjustment coefficient. At the same time, to avoid imbalance in parameter adjustment caused by excessively large or small coefficients, a range for the dynamic adjustment coefficient is set, such as 0.5 to 2.0. When the calculated coefficient exceeds this range, the boundary value of the range is automatically taken as the final coefficient. In addition, for different types of parameters in the evaluation parameter set, namely risk, sentiment, and spatial, corresponding lateral area benchmark values are set respectively. For example, the benchmark value of risk parameters is based on the historical risk parameter distribution statistics, and the benchmark value of sentiment parameters is based on the historical sentiment parameter distribution statistics, ensuring that the dynamic adjustment coefficient of each parameter conforms to its own distribution characteristics.
[0109] Step 1309 above, based on the dynamic adjustment coefficients and the initial parameters of the evaluation parameter set, completes the parameter optimization adjustment. The specific process is as follows: First, establish a one-to-one correspondence between parameters and dynamic adjustment coefficients. Initial parameters for risk categories correspond to dynamic adjustment coefficients for risk categories, initial parameters for sentiment categories correspond to dynamic adjustment coefficients for sentiment categories, and initial parameters for spatial categories correspond to dynamic adjustment coefficients for spatial categories, avoiding mismatches between coefficients and parameters. Then, perform a product operation on each initial parameter, i.e., the adjusted parameter value = initial parameter value × corresponding dynamic adjustment coefficient. This operation achieves weighted optimization of the initial parameters. For parameters with concentrated distribution (large coefficients), their adjusted values are amplified, highlighting their impact on the evaluation; for parameters with dispersed distribution (small coefficients), their adjusted values are reduced, decreasing their interference. During the operation, it is necessary to ensure that the numerical types of the initial parameter values and dynamic adjustment coefficients are consistent, both being standardized decimals, and that the operation results are retained to four decimal places to ensure the accuracy of parameter adjustment. Finally, arrange all adjusted parameter values in the order of the original evaluation parameter set to form the adjusted parameter set.
[0110] Step 1310 above, based on the adjusted parameter values and lateral area parameters, completes data fusion to form a fused parameter set. The specific process is as follows: First, the fusion method is determined, and fusion is performed by parameter item expansion. That is, based on the adjusted parameter set, a new lateral area parameter item is added, and the lateral area parameter obtained in step 1307 is used as the value of this parameter item. The lateral area parameter has been standardized in the range of 0 to 1, and the standardization method is the same as in step 1301. Then, the structure of the fused parameter set is determined. The fused parameter set is divided into three modules: the first module is the adjusted risk-related parameters, such as the adjusted risk level parameter; the second module is the adjusted emotion-related parameters, such as the adjusted emotion intensity parameter; the third module is the adjusted spatial-related parameters, such as the adjusted volume parameter; and the fourth module is the newly added lateral area parameter item. The parameters in each module are arranged in the order of the initial parameters in step 1302 to ensure that the fused parameter set contains both the optimized core evaluation parameters and the lateral area information reflecting the parameter distribution. Finally, the format of all parameters in the fused parameter set is unified, such as retaining four decimal places, to form a structurally regular and informationally complete fused parameter set.
[0111] Step 1311 above, based on the fusion parameter set, constructs a multidimensional feature vector through classification and arrangement. The specific process is as follows: First, the dimensional division of the three-dimensional structure is determined, namely, the risk dimension, the emotion dimension, and the spatial dimension. Each dimension corresponds to a type of parameter in the fusion parameter set. Then, the parameters in the fusion parameter set are classified and filtered: risk-type adjustment parameters (such as adjusted risk level parameters) are assigned to the risk dimension, emotion-type adjustment parameters (such as adjusted emotion polarity values and emotion intensity values) are assigned to the emotion dimension, and spatial-type adjustment parameters (such as adjusted volume parameters) and lateral area parameters are assigned to the spatial dimension. Next, the parameter order within each dimension is determined: the risk dimension is sorted according to the importance of risk level, such as risk level parameters taking priority; the emotion dimension is sorted according to the degree of emotion impact, such as emotion intensity parameters taking priority; and the spatial dimension is sorted according to the priority of spatial representation, such as volume parameters taking priority and lateral area parameters taking second place. Finally, the parameters of the three dimensions are concatenated sequentially to form an ordered vector structure of risk dimension parameters, emotion dimension parameters, and spatial dimension parameters. This vector is a multidimensional feature vector with a three-dimensional structure. Each element of the vector corresponds to a specific parameter of a clear dimension, ensuring that subsequent processing can accurately locate the feature information of different dimensions.
[0112] Step 1312 above, based on the multidimensional feature vector, i.e., a three-dimensional vector in the Cartesian coordinate system, containing three coordinate values x, y, and z, corresponding to risk, sentiment, and spatial dimension parameters respectively, performs spherical coordinate transformation. The specific process is as follows: First, determine the three output parameters of the spherical coordinate transformation, namely radial distance, azimuth angle, and elevation angle; calculate the radial distance, which is the vector magnitude of the multidimensional feature vector, i.e., the straight-line distance from the origin of the Cartesian coordinate system to the corresponding coordinate point of the vector. The calculation requires combining the x, y, and z coordinate values of the vector and obtaining this distance value through geometric relationships; calculate the azimuth angle, which is the angle between the projection of the multidimensional feature vector onto the xy plane and the positive x-axis. First, determine the projection coordinates (x, y) of the vector on the xy plane, then determine the quadrant of the angle based on the sign of the projection coordinates, and finally obtain the azimuth angle through angle calculation methods, with a value range of 0 to 360 degrees; calculate the elevation angle, which is the angle between the multidimensional feature vector and the positive z-axis. First, calculate the magnitude of the vector's projection onto the xy plane, i.e., the square root of x... 2 +y 2 Then, based on the proportional relationship between the projection modulus and the vector z coordinate value, the elevation angle is obtained through the angle calculation method, with a value range of 0 to 90 degrees. During the calculation process, it is necessary to ensure that the angle unit is consistent, that is, all are degrees, and the distance unit is consistent with the coordinate unit of the Cartesian coordinate system, and finally form a spherical coordinate transformation result containing radial distance, azimuth angle, and elevation angle.
[0113] Step 1313 above, using the spherical coordinate transformation result as the core input, aims to map multidimensional feature vectors to a preset unit sphere through a spherical projection algorithm, thereby achieving scale uniformity for different feature vectors and establishing a consistent coordinate benchmark for subsequent spatial distribution analysis. The specific calculation process needs to be carried out in four stages: parameter extraction and verification, implementation of projection rules, determination of projection points, and result verification, as detailed below:
[0114] First, the spherical coordinate transformation results need to be extracted and validated for accuracy. This is fundamental to ensuring the accuracy of subsequent projection calculations. The spherical coordinate transformation results contain three core parameters: radial distance, azimuth, and elevation. The radial distance refers to the straight-line distance in space from the origin to the endpoint of the multidimensional feature vector (in Cartesian coordinates) generated in step 1311. Its value must be greater than 0. If the radial distance is 0 due to the feature parameters being extremely close to zero, it should be set to a minimum value, such as 0.0001, to avoid abnormalities in subsequent normalization calculations. The azimuth refers to the projection of the multidimensional feature vector onto the xy-plane in the Cartesian coordinate system and the positive x-axis direction. The angle between the radial distance and the azimuth angle must be controlled within the range of 0 to 360 degrees. If the azimuth angle output in step 1312 exceeds this range, it needs to be cyclically corrected to conform to the specification. For example, 370 degrees is corrected to 10 degrees, and -10 degrees is corrected to 350 degrees. The elevation angle refers to the angle between the multidimensional feature vector and the positive z-axis of the Cartesian coordinate system. The value range must be controlled within the range of 0 to 90 degrees. If it exceeds this range, the difference between it and 90 degrees must be taken (e.g., 100 degrees is corrected to 80 degrees) to ensure that the parameter is valid. After the parameter extraction and correction are completed, the three parameters are temporarily stored in the order of radial distance, azimuth angle, and elevation angle as input parameters for projection calculation.
[0115] After determining the input parameters and completing the verification, it is necessary to first determine the technical definition of the preset unit sphere and the core logic of the projection rules. The preset unit sphere is a standard sphere with the origin of the Cartesian coordinate system as the center and a fixed radius of 1. The core purpose of this radius value is to unify the projection scale of all multi-dimensional feature vectors. Since the radial distance output in step 1312 will present different values due to the feature differences of different semantic paragraphs (e.g., the radial distance of paragraphs with high risk intensity may be larger, while the radial distance of paragraphs with stable emotions may be smaller), if it is directly used for subsequent spatial analysis, it is easy to cause the inability to accurately compare the feature distribution of different paragraphs due to scale differences. Therefore, the projection rules adopt a radial normalization strategy, that is, only the radial distance is scaled, while the azimuth and elevation angles remain unchanged, so as to ensure that the direction information of each feature vector is not affected after projection. The direction information is represented by the azimuth and elevation angles, and only their distances to the origin are unified, that is, all are 1, to achieve the projection goal of preserving direction and unifying scale.
[0116] Subsequently, based on the above rules, specific projection calculation operations are carried out. This process requires two steps to complete the normalization of radial distance and the determination of projection point coordinates. The first step is to calculate the normalization coefficient of the radial distance. The calculation logic for the normalization coefficient is to take the radius of a preset unit sphere (fixed at 1) as the target value and divide it by the current radial distance corrected in step 1312. For example, if the current radial distance after correction is 2, then the normalization coefficient is 1 / 2 = 0.5; if the current radial distance is 0.5, then the normalization coefficient is 1 / 0.5 = 2. During the calculation process, it is necessary to ensure that the numerical precision of the normalization coefficient is retained to four decimal places to avoid errors due to insufficient precision. The subsequent projection distance deviation; the second step is to calculate the radial distance after projection. The original radial distance corrected in step 1312 is multiplied by the normalization coefficient mentioned above. For example, when the original radial distance is 2 and the coefficient is 0.5, the product result is 1; when the original radial distance is 0.5 and the coefficient is 2, the product result is also 1. Finally, the radial distance after projection is equal to the radius of the preset unit sphere of 1. During this process, the azimuth and elevation parameters need to be confirmed simultaneously. Since the projection rules clearly retain the direction, the corrected azimuth and elevation values are directly used without any adjustment, ensuring that the directional features of the multidimensional feature vector are completely transmitted to the projection result.
[0117] After completing the radial distance normalization, it is necessary to further determine the spherical projection point of the multidimensional feature vector on the unit sphere and verify the validity of the projection result. First, the projected radial distance, corrected azimuth angle, and corrected elevation angle are combined to form the spherical coordinates of the spherical projection point, which uniquely corresponds to a spatial position on the unit sphere. To ensure that this position strictly falls on the unit sphere, a verification operation is required: the coordinates in the Cartesian coordinate system are derived from the spherical coordinates. That is, the x-coordinate is calculated by multiplying the sine of the elevation angle, the cosine of the azimuth angle, and the radial distance; the y-coordinate is calculated by multiplying the sine of the elevation angle, the sine of the azimuth angle, and the radial distance; and the z-coordinate is calculated by multiplying the cosine of the elevation angle and the radial distance. Then, the distance from the Cartesian coordinate point to the origin is calculated. If the distance is within the range of 0.9999 to 1.0001, allowing for small calculation errors, the projection point is considered valid. If it exceeds this range, the normalization coefficient calculation process needs to be checked back, and the projection operation is re-executed after eliminating coefficient calculation errors until the projection point is verified as valid.
[0118] Finally, the verified and qualified spherical projection points are the output results, and their spherical coordinates are clearly defined as (radial distance, original azimuth angle, original elevation angle). All projection points are uniformly distributed on a unit sphere with a radius of 1, which not only preserves the directional characteristics of the multidimensional feature vectors, but also eliminates the analysis interference caused by scale differences.
[0119] Step 1314 above, based on the spherical projection point (spherical coordinates), converts it back to Cartesian coordinates to determine the mapping position. The specific process is as follows: First, determine the conversion logic from spherical coordinates to Cartesian coordinates. Using the center of the unit sphere, i.e., the origin of the Cartesian coordinate system, as the reference, calculate the x, y, and z coordinate values by combining the radial distance, azimuth angle, and elevation angle of the spherical projection point; calculate the x-axis coordinate value by multiplying the sine of the elevation angle by the cosine of the azimuth angle, and then multiplying by the radial distance; calculate the y .... The x, y, and z coordinates are calculated by multiplying the sine of the azimuth angle by the radial distance. The z-axis coordinate is calculated by multiplying the cosine of the elevation angle by the radial distance. During the calculation, it is necessary to ensure that the angle parameters (azimuth and elevation angles) have been converted to the format required for angle calculation, and that the coordinate values are retained to four decimal places to ensure accuracy. The final x, y, and z coordinates are the mapped position coordinates of the spherical projection point in the Cartesian coordinate system. The point corresponding to these coordinates is located on the unit sphere and can be directly used for subsequent spatial comparison with the safe area model.
[0120] Step 1315 above involves conducting spatial relationship analysis and calculating spherical distances based on the mapped location coordinates and the preset safe area spherical model. The preset safe area spherical model is the core benchmark representing the projection range of low-risk call parameters. It needs to be preset, constructed, and trained using historical safe call data to ensure that the model parameters accurately match the characteristic patterns of safe calls made by the elderly. Then, it is integrated into the system for analysis in the current step. The specific process involves five stages: model preset, construction, training, implementation, and spatial relationship analysis, detailed below:
[0121] Before conducting spatial relationship analysis between the mapping location and the spherical model of the security area, it is necessary to first complete the pre-setting and construction of the pre-defined spherical model of the security area. The core of this model is to determine the spherical parameters that can cover the vast majority of secure call mapping locations. Specifically, the pre-setting and construction process of the model is as follows:
[0122] The first step is to prepare the data source for model building by selecting historical safe call samples of the elderly accumulated in the system. The samples must meet the risk-free judgment criteria of manual annotation, such as no risk-related keywords in the call content, stable emotional state, and no abnormal communication behavior. The sample size must be no less than 3,000 to ensure statistical validity. At the same time, it should cover safe call scenarios of different time periods and different callers. Different time periods include weekdays and holidays, and different callers include children, relatives and friends, and service organizations to avoid model bias caused by a single sample. These safe call samples are then processed one by one according to the process from steps 1301 to 1314, that is, to complete the operation of feature normalization, parameter mapping, spherical projection, etc., and finally obtain the mapped position coordinates of each safe sample on a unit sphere, forming the dataset for model building.
[0123] The second step is to determine the core parameters of the safe area spherical model: the center of the sphere and the radius R. Since the mapping position coordinates in steps 1314 are all generated based on the unit sphere constructed from the origin of the Cartesian coordinate system, the center of the safe area spherical model is preset to the origin of the Cartesian coordinate system to ensure the consistency of spatial coordinates. Then, the radius R is calculated, which is the radial distance of the mapping position coordinates of all safe samples. Since the mapping position is on the unit sphere, the radial distance is 1. Here, the mean and quantile of the original feature vector magnitude of the sample before spherical projection are actually calculated. First, the standardized values of the original feature vector magnitudes of all safe samples are counted, and then the 95th quantile of the statistical result is taken as the initial radius R. The 95th quantile is selected to cover the vast majority of safe samples while excluding extreme outliers. For example, if the value corresponding to the 95th quantile is 0.6, then the initial R value is set to 0.6.
[0124] The third step is to verify the initial radius R. 30% of the samples in the constructed dataset are randomly selected as the validation set. The percentage of samples in the validation set whose radial distance from the mapped position coordinates to the origin is less than or equal to the initial R value is counted. If the percentage is greater than or equal to 90%, it means that the initial R can effectively cover most of the safe samples and can be determined as the model radius. If the percentage is less than 90%, the quantile is adjusted (e.g., the 90th quantile is used) and the R value is recalculated and verified again until the coverage percentage of the validation set meets the standard. At this point, the initial construction of the safe area spherical model is completed. The model is defined as a sphere with the origin of the Cartesian coordinate system as the center and a radius of R (R is less than 1, such as 0.6) and its internal region.
[0125] After the model is built, its generalization ability and accuracy need to be optimized through training to avoid the model being unable to adapt to new secure call scenarios due to the limitations of historical samples. The training process is as follows: First, a batch training mechanism is adopted, adding no less than 500 recent secure call samples every month. The new samples are processed into mapped location coordinates according to the same process and added to the model training dataset. Second, the radial distance quantiles of the training dataset are recalculated. The training dataset consists of the original dataset and the new samples. If the deviation between the new quantile and the current R value is less than or equal to 5%, the current R value is maintained. If the deviation is large... If the threshold is 5%, the R-value is updated according to the new quantile to ensure that the model can adapt to subtle changes in call characteristics. The third step is to evaluate the training effect by selecting risky call samples from the same period, i.e., manually labeled as risky, and comparing their mapped location coordinates with the updated safe area spherical model. The proportion of risky samples that are misclassified as being in the safe area is calculated (misclassification rate). If the misclassification rate is less than or equal to 3%, it means that the model's ability to distinguish between safe and risky samples meets the standard. If the misclassification rate is greater than 3%, the sample labeling and data processing process is checked back, and after eliminating labeling errors or processing deviations, the model is retrained until the misclassification rate meets the requirements, thus completing the model training.
[0126] The trained safe area spherical model needs to be embedded into the system for use. The specific implementation process is as follows: the core parameters of the model (the center coordinates of the sphere are the Cartesian origin and the radius R value) are stored in the system's configuration file. The configuration file is associated with the mapping position calculation module in step 1314. When step 1314 outputs the mapping position coordinates, the system automatically reads the model parameters in the configuration file without manual intervention. At the same time, a periodic update interface for the model parameters is set up to automatically import the updated R value according to the training cycle, ensuring that the model parameters are always up-to-date.
[0127] Based on the aforementioned preset, constructed, and trained spherical model of the safe area, spatial relationship analysis and spherical distance calculation of the mapping position can be performed. The specific operation is as follows: First, determine the parameters of the currently invoked spherical model of the safe area, i.e., the center of the sphere is the origin and the radius is R. Then, calculate the spherical distance from the mapping position to the boundary of the safe area. This distance is the shortest arc length between the mapping position point on a unit sphere and the boundary sphere of the safe area (radius R). During the calculation, first determine the intersection point of the vector from the mapping position point to the origin and the boundary sphere of the safe area. The spherical coordinates (azimuth and elevation) of this intersection point are completely consistent with those of the mapping position point, only the radial distance is adjusted to R. That is, the coordinates of the intersection point in the Cartesian coordinate system can be obtained by converting the spherical coordinates and the R value. Next, calculate the distance between the mapping position point and this intersection point on a unit sphere. The arc length is calculated by first determining the angle (central angle) between the corresponding vectors of the two points (pointing from the origin to the two points), then multiplying the radian value of this angle by the radius of a unit sphere (which is 1). The resulting product is the spherical distance. Finally, the sign of the spherical distance is determined based on the relationship between the radial distance of the mapped position point and the value of R. The radial distance is the standardized radial distance calculated in step 1312: if the radial distance of the mapped position point is less than or equal to R, meaning it is inside the safe zone, the spherical distance is positive, and the magnitude of the positive value indicates the distance from the point to the boundary of the safe zone; if the radial distance of the mapped position point is greater than R, meaning it is outside the safe zone, the spherical distance is negative, and the absolute value of the negative value indicates the distance from the point to the boundary of the safe zone. This quantifies the degree of spatial deviation between the mapped position and the safe zone.
[0128] Step 1316 above, based on the spherical distance, calculates the coverage area ratio and uses it as the risk weight coefficient. The specific process is as follows: First, determine the coverage area of the spherical model of the safe area. This area is the spherical cap area corresponding to the radius R on a unit sphere, that is, the projected area of the safe area on a unit sphere. Its size is calculated based on the R value through geometric relationships. Then, determine the type of region where the mapping position is located based on the spherical distance. If the spherical distance is greater than or equal to 0, that is, the mapping position is within the safe area, then calculate the area of the sub-region where the mapping position is located, that is, the spherical cap area corresponding to the radial distance r from the origin to the mapping position. The coverage area ratio is the sub-region area / the safe area area. A larger ratio indicates that the mapped location is closer to the center of the safe area, and the lower the risk. If the spherical distance is less than 0, meaning the mapped location is outside the safe area, then the area of the sub-region outside the safe area is calculated. This is the area of the region where the mapped location is located within the remaining area after subtracting the safe area from the total area of the sphere. The coverage ratio is the area of the sub-region outside the safe area / (total area of the sphere - area of the safe area). A larger ratio indicates that the mapped location is farther away from the safe area, and the higher the risk. Finally, the coverage ratio is converted into a value between 0 and 1 as a risk weight coefficient. A larger ratio (higher risk) results in a larger coefficient, and a smaller ratio (lower risk) results in a smaller coefficient, thus achieving a quantitative representation of the risk level.
[0129] Step 1317 above identifies risk areas and determines sentiment weight coefficients based on the distribution density of mapped positions on a unit sphere. The specific process is as follows: First, the unit sphere is divided into an analysis grid, uniformly dividing it into several spherical grid cells of equal size, such as each cell having a central angle of 10 degrees × 10 degrees. Then, the number of mapped position points within each grid cell is counted, i.e., the distribution density, and a density threshold is set. Based on historical risk call data, a high density is defined as a grid cell with 5 or more points. Grid cells with densities higher than the threshold are designated as high-risk areas, and those with densities lower than the threshold are designated as low-risk areas. Then, the multidimensional features from step 1311 are combined... The emotional dimension parameters of the vector, such as the emotional polarity value, are determined as follows: if the mapping location is in a high-risk area and the emotional polarity is negative (i.e., the standardized value is less than 0.5), the emotional weight coefficient takes a higher value, such as 0.8 to 1.0; if it is in a high-risk area but the emotional polarity is positive (greater than or equal to 0.5), the coefficient takes a medium value, such as 0.5 to 0.7; if it is in a low-risk area and the emotional polarity is negative, the coefficient takes a medium-low value, such as 0.3 to 0.4; if it is in a low-risk area and the emotional polarity is positive, the coefficient takes a lower value, such as 0.1 to 0.2. The final determined value is the emotional weight coefficient, which integrates spatial distribution density and emotional state to quantify the risk correlation of the emotional dimension.
[0130] Step 1318 above first determines the weighting ratio of the risk weighting coefficient and the sentiment weighting coefficient. This ratio is based on the statistical analysis of the contribution of the two types of coefficients to risk assessment in historical risk event data. For example, the risk weighting coefficient accounts for 60% and the sentiment weighting coefficient accounts for 40%. The contribution is calculated by the success rate of accurately identifying risks using risk parameters and assisting in risk identification using sentiment parameters in historical events. Then, a weighted summation operation is performed, i.e., the comprehensive assessment score = risk weighting coefficient × 60% + sentiment weighting coefficient × 40%, and the score ranges from 0 to 1. Next, the correspondence rules between the comprehensive assessment score and the risk level are set: a score of 0 to 0.3 corresponds to a low risk level, 0.3 to 0.7 corresponds to a medium risk level, and 0.7 to 1.0 corresponds to a high risk level. Finally, the corresponding risk level is determined based on the calculated comprehensive assessment score, and this risk level is the comprehensive assessment result.
[0131] In the intelligent call assistance system for the elderly described in this embodiment of the invention, the control module 14 determines call transmission control commands based on the comprehensive evaluation results, including:
[0132] Step 1401: Compare the comprehensive evaluation result with the preset risk threshold. When the comprehensive evaluation result is less than the first threshold, it is determined to be a normal call control command.
[0133] Step 1402: When the comprehensive evaluation result is greater than the first threshold and less than the second threshold, it is determined as a reminder warning control instruction, including displaying risk warning information on the call interface;
[0134] Step 1403: When the comprehensive evaluation result is greater than the second threshold, it is determined to be a call protection control instruction, including temporarily blocking the call sound and sending a warning message to the preset guardian terminal;
[0135] Step 1404: Based on the determined call transmission control command, adjust the working status of the call transmission channel in real time to assist the elderly in the call process.
[0136] In this embodiment of the invention, when applied in a specific way, it can be implemented through the following technical solutions, for example:
[0137] Step 1401 above is based on the comprehensive evaluation results, which range from 0 to 1. The core of this evaluation is to determine normal call scenarios by comparing with a preset risk threshold. The preset risk threshold (i.e., the first threshold) needs to be calibrated using historical data to ensure it aligns with the call risk characteristics of the elderly and avoids the problem of insufficient adaptability of fixed thresholds in existing technologies. The specific preset and calculation process is as follows: First, the preset calibration of the first threshold is completed. Specifically, the following steps are taken: Select call samples of the elderly accumulated in the system over the past 6 months, with a sample size of no less than 3000. These samples must fully cover the three scenarios of manually labeled normal calls, medium-risk calls, and high-risk calls to ensure sample representativeness. Extract all normal call samples from these samples. The corresponding comprehensive evaluation results are statistically analyzed, and the 95th percentile is taken as the initial value of the first threshold. This quantile selection can cover most normal call scenarios and reduce the probability of misjudgment. Then, 30% of the samples in the historical samples are divided into a validation set, and the initial first threshold is applied to this validation set. The proportion of normal call samples in the validation set that are misjudged as abnormal calls is statistically analyzed. If the misjudgment rate is less than or equal to 3%, the initial first threshold is deemed to be suitable and is fixed as the system's preset first threshold. If the misjudgment rate is greater than 3%, the quantile is adjusted, the initial threshold is recalculated, and the verification is performed again. For example, the quantile is adjusted to the 90th percentile until the misjudgment rate meets the requirements, thus completing the preset of the first threshold.
[0138] After the above calibration process, the comprehensive evaluation result output by the optimization module 13 is then read and compared with the preset first threshold. The comparison accuracy is ensured by comparing the values bit by bit, that is, first comparing the integer part, and then comparing the values after the decimal point until the size relationship between the two is determined. When the value of the comprehensive evaluation result is less than the first threshold, the system automatically generates a normal call control command. This command will maintain the default working state of the call transmission channel, without interfering with the audio signal transmission and reception process, and without adjusting the normal display of the call interface, ensuring that the normal call experience of the elderly is not affected. At the same time, the system will synchronously store the judgment result of this normal call, the corresponding comprehensive evaluation score and the threshold comparison record in the call log, providing actual data reference for the subsequent iterative calibration of the first threshold.
[0139] In step 1402 above, a reminder and warning control instruction is generated for medium-risk scenarios. First, the second threshold is preset according to the first threshold calibration logic to form a gradient risk range. The comprehensive assessment result and the two thresholds are read, and it is determined whether the assessment result is greater than the first threshold and less than the second threshold. When both conditions are met, a reminder and warning control instruction is generated. When the instruction is executed, a pop-up risk prompt suitable for the elderly is displayed in the center of the call interface. It uses a bright but not glaring orange 24-point font, and the content focuses on risk verification prompts. It does not interfere with the transmission and reception of call audio. The prompt information continues to be displayed until the user confirms that it is closed. At the same time, the relevant time information is recorded in the call log.
[0140] In step 1403 above, a call protection control command is generated for high-risk scenarios. The comprehensive assessment result and the second threshold are read. After confirming that the assessment result is greater than the second threshold by comparing each bit, two core operations are triggered simultaneously: First, the audio control interface is called to pause the speaker output while retaining the microphone input, and a red protection instruction is displayed at the bottom of the call interface; Second, contact information is retrieved from the preset guardian information database to generate standardized warning information containing risk level, call time, and other party's number. This information is sent through both APP push and SMS backup channels to ensure that the guardian receives it. The audio blocking status is monitored in real time, and the sending and reading of warning information are recorded to form a protection closed-loop log.
[0141] Step 1404 above translates control commands into call transmission channel status adjustments, realizing the implementation of risk assessment into practical assistance. First, a mapping table is established between control commands and channel parameter adjustment strategies to determine the corresponding working states of audio transmission and reception, interface display, and external communication modules. After obtaining the determined control commands, the corresponding adjustment strategies are invoked and executed: normal commands maintain the default channel parameters; reminder commands overlay risk warnings and monitor pop-up status; protection commands adjust the audio module, interface display, and external communication module sequentially. During this process, the system reads channel parameters in real time to ensure consistency with command requirements. After the call ends, the system automatically restores the initial channel state and records the adjustment timestamp in the call log.
[0142] like Figure 2 As shown in the diagram, the real-time acquisition of user voice data stream during a call according to the present invention includes the following steps:
[0143] Real-time acquisition of user voice data stream during calls;
[0144] The speech data stream is parsed to obtain parsed data; risk keyword identification processing is performed on the parsed data to obtain risk keyword identification results; sentiment feature analysis processing is performed on the parsed data to obtain sentiment feature analysis results; a three-dimensional feature distribution model is constructed based on the risk keyword identification results and sentiment feature analysis results; the volume parameters of the three-dimensional feature distribution model in the preset feature space are calculated; the risk keyword identification results, sentiment feature analysis results, and volume parameters are combined to form multimodal feature data.
[0145] The multimodal feature data is parameterized to obtain an evaluation parameter set. A feature weight distribution cone model is constructed based on the evaluation parameter set. The lateral area parameter of the feature weight distribution cone model is calculated, and the lateral area parameter is used as a weight optimization parameter and fused into the evaluation parameter set to obtain a fused parameter set. The fused parameter set is used to synthesize feature vectors to form multidimensional feature vectors. The multidimensional feature vectors are projected to obtain the mapping positions. Based on the mapping positions, spatial distribution analysis is performed to obtain the corresponding multidimensional fusion weight coefficients. Based on the multidimensional fusion weight coefficients, the comprehensive evaluation result is calculated.
[0146] Based on the comprehensive evaluation results, the call transmission control instructions are determined.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A smart communication assistance system for the elderly, characterized in that, include: The acquisition module is used to collect the user's voice data stream during the call in real time; The processing module is used to parse the voice data stream to obtain parsed data; The parsed data is processed to identify risk keywords, and the risk keyword identification results are obtained. Sentiment feature analysis is performed on the parsed data to obtain the sentiment feature analysis results. A three-dimensional feature distribution model was constructed based on the results of risk keyword identification and sentiment feature analysis. Calculate the volume parameters of the three-dimensional feature distribution model within the preset feature space; combine the risk keyword identification results, sentiment feature analysis results, and volume parameters to form multimodal feature data; The optimization module is used to parameterize multimodal feature data to obtain an evaluation parameter set; and to construct a cone model of feature weight distribution based on the evaluation parameter set. Calculate the lateral area parameter of the cone model of feature weight distribution, and fuse the lateral area parameter as a weight optimization parameter into the evaluation parameter set to obtain the fused parameter set; synthesize the feature vectors of the fused parameter set to form a multidimensional feature vector; The multidimensional feature vectors are projected to obtain the mapping positions; based on the mapping positions, spatial distribution analysis is performed to obtain the corresponding multidimensional fusion weight coefficients. The comprehensive evaluation result is calculated based on the multi-dimensional fusion weighting coefficients; The control module is used to determine the call transmission control commands based on the comprehensive evaluation results.
2. The intelligent call assistance system for the elderly according to claim 1, characterized in that, Real-time acquisition of user voice data stream during calls, including: The system receives the user's original voice signal through a microphone array; performs noise reduction processing on the original voice signal to obtain a noise-reduced voice signal; and performs frame segmentation processing on the noise-reduced voice signal to divide the continuous voice signal into multiple voice frames. Spatial location analysis is performed on each speech frame. A three-dimensional spatial coordinate system is constructed based on the geometric layout of the microphone array. The azimuth and elevation parameters of the sound source in the spherical coordinate system are calculated to obtain a speech frame with spatial positioning information. The speech frame with spatial positioning information is pre-emphasized to obtain the pre-emphasized speech frame; the pre-emphasized speech frame is then windowed to obtain the windowed speech frame. The windowed speech frames are subjected to Fourier transform to convert them into frequency domain feature data, which is then used as the speech data stream.
3. The intelligent call assistance system for the elderly according to claim 2, characterized in that, The voice data stream is parsed to obtain parsed data; The parsed data is processed for risk keyword identification, resulting in risk keyword identification results, including: Speech recognition processing is performed on the speech data stream to convert the frequency domain feature data into a text sequence, resulting in preliminary parsed data. Semantic segmentation processing is then performed on the preliminary parsed data to divide the text sequence into multiple semantic segments based on semantic coherence, resulting in structured parsed data. The structured parsed data is segmented to obtain a keyword set; the keyword set is then matched and compared with a preset risk keyword database to identify the successfully matched risk keywords. The risk keywords are spatially mapped, and each risk keyword is mapped to a preset two-dimensional semantic plane according to its semantic features, forming a spatial point set of risk keywords; Based on the spatial point set of risk keywords, semantic association triangles between adjacent risk keywords are constructed, the area value of each semantic association triangle is calculated, and the semantic association strength between risk keywords is determined according to the size of the area value. The validity of risk keywords is verified based on the semantic association strength and semantic paragraph context information to obtain valid risk keywords. The valid risk keywords are then classified and labeled, and the risk level labels are adjusted according to the semantic association strength to obtain the risk keyword identification results containing the risk level labels.
4. The intelligent call assistance system for the elderly according to claim 3, characterized in that, Sentiment feature analysis is performed on the parsed data to obtain the sentiment feature analysis results. A three-dimensional feature distribution model is constructed based on the results of risk keyword identification and sentiment feature analysis, including: Sentiment features are extracted from structured parsed data. By analyzing the density of sentiment words and the trend of sentiment intensity in semantic paragraphs, sentiment feature analysis results containing sentiment polarity and sentiment intensity values are obtained. The risk level identifiers in the risk keyword identification results are converted into risk intensity values, and the sentiment polarity values and sentiment intensity values in the sentiment feature analysis results are used as the first and second dimension parameters, respectively. Based on the risk intensity value, sentiment polarity value, and sentiment intensity value, each semantic paragraph is mapped to a three-dimensional coordinate system to form a three-dimensional spatial point set; Based on the distribution density of the three-dimensional spatial point set, a minimum convex polyhedron enclosing all spatial points is constructed, and the minimum convex polyhedron is used as a three-dimensional feature distribution model.
5. The intelligent call assistance system for the elderly according to claim 4, characterized in that, Calculate the volume parameters of the three-dimensional feature distribution model within a preset feature space; The results of risk keyword identification, sentiment feature analysis, and volume parameters are combined to form multimodal feature data, including: The smallest convex polyhedron is orthogonally projected along a direction perpendicular to the bottom surface of the preset feature space to obtain the projected contour; the boundary extraction process is performed on the projected contour to identify the outer boundary curve of the projected contour. Based on the outer boundary curve, calculate the minimum circumcircle of the projected profile, and use the radius of the minimum circumcircle as the radius parameter of the cylinder base; use the maximum height difference of the minimum convex polyhedron in the vertical direction as the cylinder height parameter. Based on the radius parameter of the cylinder base and the height parameter of the cylinder, calculate the volume parameter occupied by the three-dimensional feature distribution model in the preset feature space, where the volume parameter is equal to the product of the area of the base circle and the height; The risk keyword identification results, sentiment feature analysis results, and volume parameters are synchronized and aligned according to timestamps, and encapsulated into a unified data structure to jointly constitute multimodal feature data.
6. The intelligent call assistance system for the elderly according to claim 5, characterized in that, The multimodal feature data is parameterized to obtain an evaluation parameter set; a cone model of feature weight distribution is constructed based on the evaluation parameter set, including: The multimodal feature data is normalized by converting the risk level identifier in the risk keyword identification results, the sentiment polarity value and sentiment intensity value in the sentiment feature analysis results, and the volume parameter into standardized values to obtain a set of normalized feature values. Based on the normalized eigenvalue set, the main feature components are extracted and used as the initial parameters of the evaluation parameter set. Based on the initial parameters of the evaluation parameter set, a three-dimensional parameter distribution space is constructed, and each evaluation parameter set is mapped to the corresponding coordinate point in the three-dimensional parameter distribution space; Based on the distribution characteristics of all coordinate points in the three-dimensional parameter distribution space, the center point and boundary range of the evaluation parameter set distribution are identified; A cone model of feature weight distribution is constructed by taking the center point of the evaluation parameter set distribution as the vertex of the cone and the boundary range as the bottom contour of the cone.
7. The intelligent call assistance system for the elderly according to claim 6, characterized in that, Calculate the lateral area parameter of the feature weight distribution cone model, and fuse the lateral area parameter as a weight optimization parameter into the evaluation parameter set to obtain the fused parameter set, including: Geometric parameters are extracted from the feature weight distribution cone model to obtain the cone height and base radius parameters; Based on the cone height parameter and the base radius parameter, the lateral area parameter of the feature weight distribution cone model is calculated, where the lateral area parameter reflects the coverage and concentration of the parameter distribution; Based on the lateral area parameter, the dynamic adjustment coefficients of each evaluation parameter are determined, wherein the dynamic adjustment coefficients are inversely proportional to the lateral area parameter; The adjusted parameter values are obtained by multiplying the dynamic adjustment coefficient with the corresponding parameter in the evaluation parameter set. The adjusted parameter values are then fused with the lateral area parameter to obtain a fused parameter set.
8. The intelligent call assistance system for the elderly according to claim 7, characterized in that, The fusion parameter set is used to synthesize feature vectors to form multidimensional feature vectors; The multidimensional feature vector is projected to obtain the mapped position, including: The parameters in the fusion parameter set are classified and arranged according to the risk dimension, emotional dimension and spatial dimension to form a multidimensional feature vector with a three-dimensional structure; A spherical coordinate transformation is performed on the multidimensional feature vectors to convert the multidimensional feature vectors in the Cartesian coordinate system into spherical coordinate transformation results. The spherical coordinate transformation results include radial distance, azimuth angle, and elevation angle parameters in the spherical coordinate system. Based on the results of spherical coordinate transformation, the multidimensional feature vector is projected onto a preset unit sphere to obtain spherical projection points; Transform the spherical coordinates of the spherical projection point back to Cartesian coordinates to obtain the mapped position coordinates on the unit sphere, which is then used as the mapped position.
9. The intelligent call assistance system for the elderly according to claim 8, characterized in that, Based on the mapping location, spatial distribution analysis is performed to obtain the corresponding multidimensional fusion weight coefficients; Based on the multidimensional fusion weighting coefficients, the comprehensive evaluation results are calculated, including: The spatial relationship between the mapped location and the preset safe area spherical model is analyzed, and the spherical distance from the mapped location to the boundary of the safe area is calculated. Based on the spherical distance, the coverage area ratio of the region where the mapping location is located is calculated, and the coverage area ratio is used as the risk weight coefficient. Based on the distribution density of the mapped positions on the sphere, high-risk and low-risk areas are identified, and the sentiment weight coefficient is obtained. The comprehensive assessment result is calculated based on the risk weighting coefficient and the sentiment weighting coefficient, and the comprehensive assessment result reflects the risk level of the call.
10. The intelligent call assistance system for the elderly according to claim 9, characterized in that, Based on the comprehensive evaluation results, the call transmission control instructions are determined, including: The comprehensive assessment result is compared with the preset risk threshold. When the comprehensive assessment result is less than the first threshold, it is determined to be a normal call control command. When the comprehensive assessment result is greater than the first threshold and less than the second threshold, it is determined to be a reminder warning control instruction, including displaying risk prompt information on the call interface; When the comprehensive evaluation result exceeds the second threshold, it is determined to be a call protection control instruction, including temporarily blocking the call sound and sending a warning message to the preset guardian terminal; Based on the determined call transmission control instructions, the working status of the call transmission channel is adjusted in real time to assist the elderly in their call process.
Citation Information
Patent Citations
CNN-based call quality inspection method, apparatus and device, and storage medium
CN121056563A
Coordinating voice calls between representatives and customers to influence an outcome of the call
US20170187880A1