Charging control method based on voice interaction, charging control system and storage medium

By obtaining and analyzing user voice data, charging authentication and safety authentication are performed, the problems of complex operation of charging piles and poor coordination of multiple ends are solved, convenient and safe charging control is achieved, and user experience and system security are improved.

CN120396746APending Publication Date: 2025-08-01AUTEL UNITED CREATION SOFTWARE DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510606589.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing charging piles have complex operation and poor user experience, especially for the elderly or users who are not familiar with the operation, and have poor multi-terminal coordination and lack of voice interaction coordination and scene adaptability, resulting in inconvenience in use and insufficient security.

Method used

By obtaining user voice data, performing fusion analysis and processing, judging charging authentication and performing safety authentication, charging control is realized, combining voice recognition, scene recognition and multi-terminal collaboration to optimize the interaction method.

Benefits of technology

Simplify the operation process, improve the charging success rate and efficiency, provide convenient and safe charging methods, ensure timely data synchronization between devices, and enhance system security and consistent user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120396746A_ABST
    Figure CN120396746A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of charging control, in particular to a voice interaction-based charging control method, a voice interaction-based charging control system and a storage medium, and the voice interaction-based charging control method comprises the steps of obtaining first voice data of a user; performing fusion analysis processing on the first voice data to obtain target voice information; judging whether to perform charging authentication according to the target voice information; if yes, performing security authentication on the user to obtain an authentication result corresponding to the user; and performing charging control according to the authentication result and the target voice information. According to the method, voice interaction and fusion analysis processing are introduced, so that the operation process of charging control is optimized, the user experience is improved, and the safety is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of charging control, and particularly to a charging control method, a charging control system and a storage medium based on voice interaction. Background Art

[0002] With the popularization of new energy vehicles, the usage frequency of charging piles is continuously increasing. In the prior art, when using a charging pile, users mainly rely on a mobile phone APP or the screen of the charging pile for operation. This operation method has many problems.

[0003] First of all, users need to click multiple times to complete the charging process. Although this method is common, it has the problems of complex operation steps and easy errors. Especially for the elderly or users who are not familiar with the operation, it is difficult to use. In addition, there are also problems with unnatural interaction in the prior art. For example, when driving or carrying items, it is inconvenient to manually operate the charging pile, especially when operating in bad weather or at night, the user experience is affected.

[0004] In addition, the multi-terminal collaboration between the charging pile, the mobile phone APP and the smart speaker is poor, and there is a lack of an effective voice interaction collaboration mechanism, resulting in untimely data synchronization between devices and a fragmented user experience. The existing system also has deficiencies in scene adaptability. The voice interaction requirements are different in different scenarios, and the prior art lacks the ability of scene adaptation and cannot automatically adjust the interaction method according to the environment. These problems limit the convenient use and good experience of users in different environments. Summary of the Invention

[0005] An object of an embodiment of the present invention is to provide a charging control method, a charging control system and a storage medium based on voice interaction, which are used to solve the technical problems of inconvenient operation, unnatural interaction, inability to flexibly collaborate among multiple terminals and non-adaptability to scenarios in the traditional charging pile scenario.

[0006] In a first aspect, an embodiment of the present invention provides a charging control method based on voice interaction, which is applied to a charging control system. The method includes:

[0007] Obtain the first voice data of the user;

[0008] Perform fusion analysis processing on the first voice data to obtain target voice information;

[0009] Judge whether to perform charging authentication according to the target voice information;

[0010] If so, perform security authentication on the user to obtain the authentication result corresponding to the user;

[0011] Perform charging control according to the authentication result and the target voice information.

[0012] In a second aspect, a charging control system is provided. The charging control system includes a memory and a processor. The memory is connected to the processor. The processor is configured to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the charging control system implements the voice interaction-based charging control method as described in the first aspect.

[0013] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the voice interaction-based charging control method as described in the first aspect.

[0014] In the embodiments implemented by the above voice interaction-based charging control method, charging control system, and computer-readable storage medium, first, the first voice data of the user is obtained. Secondly, the first voice data is subjected to fusion analysis processing to obtain target voice information. Then, it is determined whether to perform charging authentication according to the target voice information. If so, security authentication is performed on the user to obtain an authentication result corresponding to the user. Finally, charging control is performed according to the authentication result and the target voice information. In this embodiment, by obtaining the first voice data of the user and performing fusion analysis, the user can complete the charging operation through voice commands, reducing the dependence on the mobile phone APP or the charging pile screen, realizing effective voice interaction and coordination among the charging pile, mobile phone APP, and smart speaker, ensuring timely data synchronization between devices, avoiding the problem of information silos, simplifying the operation process. The voice interaction method enables the user to charge without manual operation when driving or carrying items. Especially in situations such as driving, carrying items, bad weather, or at night when manual operation is inconvenient, voice control provides a more convenient and safe usage method. Further, before performing charging authentication, user security authentication is performed through the target voice information to ensure that only authenticated users can perform charging operations, enhancing the security of the system. Through the fusion analysis of voice data, the interaction method can be adaptively adjusted in different scenarios, enhancing the coordination among the charging pile, mobile phone APP, and smart speaker, and providing a more consistent and seamless user experience. Therefore, in this embodiment, through the automated voice recognition and authentication process, errors caused by the complexity of manual operation can be reduced, and the success rate and efficiency of the charging operation can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 is a schematic structural diagram of a charging control system in an embodiment of the present invention;

[0017] Figure 2 is a schematic flowchart of a charging control method based on voice interaction in an embodiment of the present invention;

[0018] Figure 3 is an implementation flowchart of a fusion analysis and processing in an embodiment of the present invention;

[0019] Figure 4 is a schematic structural diagram of a charging control device based on voice interaction in an embodiment of the present invention. Detailed implementation manners

[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0021] It should be noted that if there is no conflict, the various features in the embodiments of the present invention can be combined with each other, and all are within the protection scope of the present invention. In addition, although functional module division is carried out in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. Furthermore, the terms "first", "second", "third", etc. used in the present invention do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.

[0022] Please refer to Figure 1 , Figure 1 is a schematic structural diagram of a charging control system. In Figure 1 , the charging control system 10 adopts a hierarchical architecture design, including a user terminal layer 20, a voice interaction layer 30, a service processing layer 40, and a cloud service layer 50.

[0023] Among them, the user terminal layer 20 includes a charging pile terminal, a mobile phone APP terminal, and a smart speaker terminal, which are used to receive user voice instructions and provide voice feedback.

[0024] Among them, the voice interaction layer 30 includes a voice recognition module, a voice synthesis module, a local voice processing module, and a scene recognition module, which realizes the recognition, processing, and scene adaptation of voice instructions.

[0025] Among them, the service processing layer 40 includes an identity authentication module, a payment processing module, a charging control module, and a status management module, which are responsible for user authentication, payment processing, charging control, and status management.

[0026] Among them, the cloud service layer 50 includes a data synchronization service, a scene management service, a user preference service, and a security authentication service, which realizes multi-terminal data synchronization, scene management, user preference learning, and security authentication.

[0027] Among them, each layer communicates through a standard interface to ensure the scalability and maintainability of the system. Through this layered architecture design, the present invention realizes the intelligent charging control functions of multi-terminal collaboration, scene adaptation, and security and reliability.

[0028] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of a charging control method based on voice interaction provided by an embodiment of the present invention, and is applied to a charging control system. The method includes the following steps:

[0029] S10. Obtain the first voice data of the user.

[0030] Among them, the above-mentioned first voice data refers to the audio information transmitted by the user through a voice input device (such as a charging pile terminal, a mobile phone APP terminal, and a smart speaker terminal, etc.). This audio information may contain the original voice signal of its charging instruction or related request. It can be what the user says to terminal devices such as a charging pile, a mobile phone APP, or a smart speaker, such as "start charging", "fully charge", "query power", etc.

[0031] Among them, the user refers to an individual who is using or intends to use a charging pile for charging.

[0032] Among them, the obtaining process can be that the voice input device (such as a charging pile terminal, a mobile phone APP terminal, and a smart speaker terminal, etc.) continuously listens to the surrounding sounds.

[0033] It can be seen that this embodiment can effectively obtain the first voice data of the user, and further provide a basis for subsequent charging control and user interaction.

[0034] S20. Perform fusion analysis processing on the first voice data to obtain target voice information.

[0035] Among them, the fusion analysis process refers to comprehensively using multiple technologies and data sources to analyze voice data in order to extract meaningful information. The fusion analysis process may include, but is not limited to, noise reduction, speech recognition, natural language processing, context understanding, etc.

[0036] Among them, the target voice information refers to the instructions or information that can accurately reflect the user's intention and are extracted from the original voice data after analysis and processing.

[0037] For example, when the user says "Start charging", the target voice information is the instruction "Start charging"; when the user says "Charge my car to 80%", the target voice information is "Charge the vehicle battery to 80%".

[0038] Specifically, the above-mentioned fusion analysis process can refer to the detailed description in S201 - S204, and will not be repeatedly defined here.

[0039] It can be seen that in this embodiment, the target voice information can be effectively extracted from the first voice data to ensure that the user's needs are accurately responded to.

[0040] S30. Determine whether to perform a charging authentication according to the target voice information.

[0041] Among them, the charging authentication refers to the process of verifying the user's identity and permissions before performing a charging operation to verify whether the user's identity is legal and whether the user has the permission to use the charging pile for charging.

[0042] In a specific implementation, extract the user's charging request from the target voice information and determine whether authentication is required. For example, determine whether the user requests to start, stop, or adjust charging; according to the type of the user's request and the security policy set by the system, determine whether identity verification is required. Usually, requests involving payment or sensitive operations require authentication.

[0043] It can be seen that in this embodiment, it can be effectively determined whether to perform a charging authentication to ensure the safety and compliance of the charging operation.

[0044] S40. If so, perform a security authentication on the user to obtain the authentication result corresponding to the user.

[0045] Among them, the security authentication refers to the process of confirming the user's identity and permissions through specific verification methods to ensure the safety of the charging operation and prevent unauthorized access.

[0046] Among them, the authentication result is the final output of the authentication process, which may include, but is not limited to, the status of success or failure, user identity information, permission level, and possible reasons or additional information.

[0047] Among them, in the specific implementation process of performing security authentication on the user, the system prompts the user to input necessary authentication information, such as account number, password, fingerprint, facial recognition data, etc. through the user terminal layer (such as the charging pile screen, mobile phone APP); the collected authentication information is transmitted to the service processing layer or the cloud service layer through a secure channel for verification; after receiving the authentication information, the security authentication module in the service processing layer or the cloud service layer executes the predetermined authentication logic, including: checking whether the identity information such as the account number and password input by the user is consistent with the information stored in the system; checking whether the user has the permission to use the current charging pile for charging; detecting the security of the user device and detecting whether there is any abnormal login behavior, etc. After completing the authentication logic, the system generates an authentication result and returns it to the service processing layer, and then the service processing layer decides the subsequent operations.

[0048] It can be seen that through security authentication in this embodiment, unauthorized access and illegal use can be effectively prevented, ensuring the security and legality of user requests.

[0049] S50. Perform charging control according to the authentication result and the target voice information.

[0050] Among them, charging control is a process in which the system actually controls the charging pile to charge according to the authentication result and the specific needs of the user.

[0051] Among them, if the authentication is passed and the target voice information is "start charging", then the charging operation is executed; if the authentication fails or the target voice information does not meet the expectation, the charging request is rejected or the user is prompted to try again.

[0052] Among them, when it is determined to start charging, the charging control system will issue specific charging parameters (such as voltage, current, etc.) to the charging pile to start the actual charging process.

[0053] Specifically, if the authentication result is passed, according to the charging information in the target voice information, a charging control instruction is generated; the charging control instruction is parsed to obtain the charging parameters and the charging operation is executed.

[0054] Furthermore, during the charging process, the charging control system continuously updates the status and synchronizes the status to the cloud service layer. The user terminal layer periodically feeds back the charging status to the user. After the charging is completed, the charging control system generates a bill and sends a payment request to the cloud service layer. The cloud service layer processes the payment request, returns the payment result after the payment is completed, and the user terminal layer displays the payment confirmation information to the user. After the whole process is completed, the user terminal layer displays a charging completion prompt to the user.

[0055] Specifically, after receiving a charging request, the charging control system first verifies the charging quantity in the charging request to ensure that the charging process meets safety requirements, including voltage inspection to prevent overvoltage or undervoltage situations, current limitation to avoid overload risks, temperature monitoring to ensure that overheating does not occur during charging, and setting safety thresholds to ensure that the charging system operates within a safe range. Further, by analyzing the load, the current power load situation is evaluated, and combined with an efficiency optimization strategy, the energy efficiency of the charging process is improved. At the same time, the system calculates the cost, optimizes the charging fee, and provides a time estimate to predict the charging completion time to improve the user experience. Further, based on the calculation results, the system adjusts the power, dynamically adjusting the charging power according to real-time requirements. At the same time, temperature control ensures that the device does not overheat, current limitation avoids battery overload, and voltage stability maintains the safety and consistency of the charging process. During the charging process, the system continuously monitors the status and analyzes the charging status in real time. If an abnormal situation is detected, the system immediately triggers an exception handling mechanism, including adjusting the charging strategy, issuing an alarm, or aborting the charging to ensure the safe and stable operation of the system.

[0056] It can be seen that in this embodiment, effective charging control can be performed according to the authentication result and the target voice information, ensuring that the user's request is executed safely and accurately.

[0057] In this embodiment, by obtaining the user's first voice data and performing fusion analysis, the user can complete the charging operation through voice commands, reducing the dependence on the mobile phone APP or the charging pile screen, realizing effective voice interaction and coordination among the charging pile, the mobile phone APP, and the smart speaker, ensuring timely data synchronization among devices, avoiding the problem of information islands, simplifying the operation process. The voice interaction method enables the user to charge without manual operation when driving or carrying items. Especially in situations where manual operation is inconvenient, such as driving, carrying items, bad weather, or at night, voice control provides a more convenient and safe way of use; further, before charging authentication, user safety authentication is performed through the target voice information to ensure that only authenticated users can perform charging operations, improving the security of the system; through the fusion analysis of voice data, the interaction method can be adaptively adjusted in different scenarios, enhancing the coordination among the charging pile, the mobile phone APP, and the smart speaker, providing a more consistent and seamless user experience. Therefore, through the automated voice recognition and authentication process in this embodiment, errors caused by the complexity of manual operation can be reduced, and the success rate and efficiency of the charging operation are improved.

[0058] S201. In one embodiment, the fusion analysis and processing of the first voice data to obtain the target voice information includes: performing voice recognition on the first voice data to obtain the first voice information; obtaining at least one set of fusion parameters; and performing fusion analysis and processing on the first voice information according to the at least one set of fusion parameters to obtain the target voice information.

[0059] Among them, speech recognition is the process of converting speech signals into text information, which involves steps such as speech signal preprocessing, feature extraction, acoustic models, and decoding.

[0060] Among them, the fusion parameters are various data sources and information used to optimize and adjust the speech recognition results.

[0061] Specifically, at least one set of fusion parameters can include, but is not limited to, collecting environmental data from multiple sources, which can include, but is not limited to, sensor data, time information, user behavior, and environmental parameters.

[0062] Specifically, the sensor data obtains information such as the temperature, humidity, and light of the current environment through built-in or external sensors. For example, if the temperature is very high, it may affect the charging efficiency of the battery.

[0063] Specifically, the time information can be information such as the current date, time, time period (such as morning, afternoon, evening), and whether it is a holiday. For example, the utilization rate of charging piles may be higher during holidays.

[0064] Specifically, the user behavior is to analyze the user's past usage data to understand the user's charging habits, such as the time period when the user often charges and the charging piles the user prefers to use. For example, the user often charges at a charging pile near home in the evening.

[0065] Specifically, the environmental parameters are to obtain the user's current geographical location information and the surrounding noise level. For example, if the user is in a noisy environment, the system may adopt a more accurate speech recognition model.

[0066] Among them, the fusion analysis and processing refers to the process of using the fusion parameters to deeply analyze the preliminary speech information, combining the context, environmental factors, etc., to more accurately understand the user's intention.

[0067] Optionally, in the specific implementation of performing fusion analysis processing on the first voice information according to the at least one set of fusion parameters to obtain target voice information, the true intention of the user is analyzed by combining the first voice information and the at least one set of fusion parameters. For example, when the user says "Start charging", combined with the current time and the user's location, the system can infer that the user hopes to "start charging at the nearest available charging pile"; using natural language processing technology, classify and identify the user's intention. For example, identify "Start charging" as an "electricity charging request" intention; extract key information from the first voice information, such as the charging amount, charging time, etc. For example, if the user says "Charge to 80%", then extract the charging amount as 80%; further, the system can have multiple rounds of conversations with the user to obtain more information or confirm the user's intention. For example, if the system is not sure about the specific needs of the user, it can ask "How much do you need to charge?".

[0068] For example, when the user says "Start charging", the system combines the current time of 3 pm, the user is near home, and the charging pile is available, so it understands the target voice information as "Start charging at the charging pile near home, using the default charging amount".

[0069] It can be seen that in this embodiment, the first voice information can be effectively subjected to fusion analysis processing to ensure the speech recognition accuracy and adaptability under different environments and conditions.

[0070] S202. In one embodiment, the performing speech recognition on the first voice data to obtain first voice information includes: performing data preprocessing on the first voice data to obtain the preprocessed first voice data; performing feature extraction on the preprocessed first voice data to obtain first feature parameters; inputting the first feature parameters into a preset acoustic model for processing to obtain second voice data output by the preset acoustic model; and inputting the second voice data into a preset language model for processing to obtain the first voice information output by the preset language model.

[0071] Among them, data preprocessing is to perform preliminary processing on the original voice data to remove interference and highlight effective information, preparing for subsequent processing.

[0072] Among them, data preprocessing may include but is not limited to noise reduction processing, endpoint detection, frame segmentation processing, and pre-emphasis. Specifically, noise reduction processing uses tools such as filters to remove background noise, such as wind noise and traffic flow noise, to remove background interference and improve the clarity of the voice signal. Endpoint detection is used to identify the start and end points of the voice, exclude the silent segments, and obtain effective voice segments. Frame segmentation processing divides the voice data into continuous small segments to ensure the time continuity of the voice data. Pre-emphasis enhances the high-frequency signal, compensates for the high-frequency attenuation of the voice signal, and improves the stability of feature extraction.

[0073] Among them, feature extraction refers to extracting key features, that is, the first feature parameters, from the preprocessed first speech data, and these key features can represent the main information of the speech.

[0074] Among them, the first feature parameters refer to the parameters obtained through feature extraction, such as MFCC, spectral features, energy features, and fundamental frequency features, etc.

[0075] Among them, the feature extraction process of MFCC features (Mel Frequency Cepstral Coefficients) is to calculate the Mel Frequency Cepstral Coefficients, simulating the auditory characteristics of the human ear. Spectral features are used to analyze the energy distribution of different frequencies. Energy features are used to measure the speech intensity, such as the volume size. Fundamental frequency features are used to assist in judging the speech pitch information, such as high and low tones.

[0076] Among them, the preset acoustic model refers to a pre-trained acoustic model, which is used to convert the feature parameters into speech data.

[0077] Specifically, the process of inputting the first feature parameters into the preset acoustic model for processing is to input the first feature parameters into the Transformer encoder; use the attention mechanism to focus on key information, improve the understanding ability of long-sequence speech, optimize the model output through the CTC loss, achieve effective alignment of variable-length speech sequences, perform model quantization processing, improve the calculation efficiency, and adapt to low-power devices; output the second speech data.

[0078] Among them, the second speech data is the output after being processed by the preset acoustic model, and it is closer to the representation of real speech.

[0079] Among them, the preset language model is a pre-trained language model, which is used to convert speech data into text information.

[0080] Specifically, in the process of inputting the second speech data into the preset language model for processing, input the second speech data into the Transformer decoder; adopt beam search to improve the decoding accuracy and find the most likely text sequence; combine language constraints to ensure that the recognition results conform to grammar rules; enhance semantic coherence through context modeling; output the first speech information.

[0081] It can be seen that in this embodiment, the first speech data can be effectively processed, and finally accurate language information is generated to execute the user request.

[0082] S203. In one embodiment, performing fusion analysis processing on the first voice information according to the at least one set of fusion parameters to obtain target voice information includes: parsing the at least one set of fusion parameters to obtain the parameter type of each set of fusion parameters; determining a fusion mode according to the parameter type; performing fusion processing on the at least one set of fusion parameters according to the fusion mode to obtain first fusion data; extracting feature parameters from the first fusion data to obtain second feature parameters; inputting the second feature parameters into a preset classification model for scene recognition to obtain a first scene label output by the preset classification model; and analyzing and processing the first voice information according to the first scene label to obtain target voice information.

[0083] Among them, an implementation process of a fusion analysis processing of S203 is as Figure 3 shown.

[0084] Among them, the parameter type refers to the specific category of the fusion parameter, such as environmental data, time information, user behavior, etc.

[0085] Among them, the fusion mode refers to the way of performing fusion processing on the fusion parameters determined according to the parameter type. For example, weighted average, feature-level fusion, decision-level fusion, etc.

[0086] Among them, the first fusion data refers to the data after fusion processing, which synthesizes information from multiple sources.

[0087] Among them, the second feature parameter refers to the feature parameter extracted from the first fusion data, such as time series feature, spatial feature, behavior feature, environmental feature, etc.

[0088] Among them, the preset classification model refers to a machine learning model pre-trained for scene recognition, such as support vector machine (SVM), random forest, neural network, etc.

[0089] Among them, the first scene label refers to the scene recognition result output by the preset classification model, which is used to identify the environmental category where the user is currently located, such as "home", "office", "public place", etc.

[0090] Specifically, in the process of parsing the at least one set of fusion parameters to obtain the parameter type of each set of fusion parameters, temperature and humidity are parsed from sensor data; time period and holiday are parsed from time information; device usage habits and interaction patterns are parsed from user behavior data; geographical location and noise level are parsed from environmental parameters; and the parsed parameters are classified. For example, temperature and humidity are classified as environmental data type, time period and holiday are classified as time information type, device usage habits and interaction patterns are classified as user behavior type, and geographical location and noise level are classified as environmental parameter type.

[0091] Among them, in the process of determining the fusion mode according to the parameter type, for numerical parameters (such as temperature, humidity), weighted average can be used for fusion; for categorical parameters (such as time period, holiday), feature-level fusion can be used; for parameters from different sources, decision-level fusion can be used, and there is no unique limitation here.

[0092] For example, assume that the fusion parameters include temperature (25°C), humidity (60%), time period (morning), geographical location (office), and noise level (low). After parsing, it is determined that temperature and humidity belong to the environmental data type, the time period belongs to the time information type, and geographical location and noise level belong to the environmental parameter type. According to the parameter type, weighted average is selected to fuse temperature and humidity, and feature-level fusion is selected to fuse the time period, geographical location, and noise level.

[0093] Among them, the process of fusion processing is to perform fusion processing on at least one set of fusion parameters according to the determined fusion mode to obtain the first fusion data. The purpose of fusion is to integrate information from multiple sources to form a more comprehensive and accurate data representation.

[0094] For example, according to the fusion mode determined in the previous step, temperature (25°C) and humidity (60%) are weighted averaged to obtain a comprehensive environmental comfort index; the time period (morning), geographical location (office), and noise level (low) are feature-level fused to obtain a comprehensive scene description index. Finally, these two indexes constitute the first fusion data.

[0095] Among them, the process of feature extraction is to extract features from the first fusion data, and the extracted features include temporal features (data change trend), spatial features (association between geographical location and scene), behavioral features (user habit pattern), and environmental features (temperature, humidity, light, etc.). This process improves the data representation ability and enables the system to accurately describe the scene where the user is currently located.

[0096] For example, features are extracted from the comprehensive environmental comfort index and the comprehensive scene description index in the first fusion data, and the obtained features include temporal features (such as "temperature rises in the morning"), spatial features (such as "office"), behavioral features (such as "charging during weekdays during the day"), and environmental features (such as "comfortable environment"). These features constitute the second feature parameter.

[0097] Among them, the process of scene recognition is to organize the second feature parameters into a feature vector as the input of the classification model; use a pre-trained classification model (such as SVM, random forest, neural network, etc.) to classify the feature vector and calculate the confidence of each category; generate a scene label according to the classification result to identify the environmental category where the current user is located. For example, "home", "office", "public place", etc.

[0098] For example, organize the second feature parameters (temporal feature, spatial feature, behavior feature, environmental feature) into a feature vector and input it into a preset classification model (such as SVM). The classification result output by the classification model is "office" and the confidence is 95%. Therefore, the first scene label is "office".

[0099] Among them, the purpose of analysis and processing is to more accurately understand the true intention of the user according to the scene where the user is located.

[0100] Specifically, the first scene label provides important scene information, which can help the system better understand the user's voice command. For example, in the "office" scene, when the user says "Please help me charge", it may mean "Use the fastest charging method"; while in the "home" scene, when the user says "Please help me charge", it may mean "Use the cheapest charging method".

[0101] Optionally, perform status update based on the classification result of the scene label; determine whether it is necessary to switch to a new scene mode; once the switching condition is met, perform scene switching and adjust the system configuration to adapt to the new environmental requirements.

[0102] It can be seen that in this embodiment, through multi-source information fusion of scene recognition and speech understanding, it is possible to effectively identify the scene where the user is located, more accurately understand the true intention of the user, thereby providing more intelligent and personalized services, and performing adaptive adjustments to improve the user experience.

[0103] S204. In one embodiment, the analyzing and processing the first voice information according to the first scene label to obtain target voice information includes: performing scene recognition on the first voice information to obtain a second scene label; determining whether the first scene label and the second scene label are within the same label range; if so, determining that the first voice information is the target voice information; or, if not, performing prediction processing on the first voice information according to the first scene label to obtain the target voice information.

[0104] Among them, scene recognition is to use a preset classification model or algorithm to identify the environment or usage scenario where the voice information is located and generate a corresponding scene label.

[0105] Specifically, during the process of scene recognition of the first voice message, the first voice message is input into a scene recognition model. The model analyzes keywords, context, etc. in the voice message, matches them with the preset scene classification criteria, and outputs the corresponding second scene label.

[0106] Among them, the second scene label is the scene classification result obtained by performing scene recognition on the first voice message, and is used to represent the specific environment or situation where the user is located.

[0107] Among them, the label range refers to the classification category or level of the scene label, and is used to compare whether two scene labels belong to the same category or level.

[0108] Specifically, during the process of determining whether the first scene label and the second scene label are within the same label range, the classification categories or levels of the two scene labels are compared.

[0109] Further, when the two scene labels are within the same label range, the system considers the preliminary judgment to be accurate and directly uses the first voice message as the target voice message.

[0110] Further, when the two scene labels are not within the same label range, the system reinterprets and predicts the first voice message according to the first scene label to more accurately reflect the user's intention.

[0111] Among them, the prediction process is to speculate and adjust the voice message according to a specific scene label to generate an accurate target voice message.

[0112] It can be seen that in this embodiment, by comparing the results of two scene recognitions, it is possible to more accurately determine the environment or situation where the user is located, thereby more accurately understanding the user's voice command, and further improving the accuracy and adaptability of voice recognition, and being able to better meet the user's needs in different scenarios.

[0113] In one embodiment, the determining whether to perform charging authentication according to the target voice message includes: extracting multiple keywords from the target voice message; determining whether the multiple keywords exist in a preset instruction database; if so, determining to perform charging authentication; or, if not, determining not to perform charging authentication.

[0114] Among them, a keyword refers to a keyword that can represent the user's intention or requirement in the target voice message. For example, "charging", "start", "stop", etc.

[0115] Among them, the preset instruction database refers to a set of instruction vocabulary libraries preset by the system and used to store legal instructions related to charging operations. For example, it includes instructions such as "start charging", "stop charging", "charging mode", etc.

[0116] Specifically, during the process of extracting multiple keywords from the target voice information, natural language processing operations such as word segmentation and part-of-speech tagging are performed on the target voice information to extract keyword vocabularies that can represent the user's intention from the target voice information.

[0117] Specifically, during the process of determining whether the multiple keywords exist in the preset instruction database, each extracted keyword is compared with the instructions in the preset instruction database to check for matching items.

[0118] Furthermore, when the extracted keyword matches the instruction in the preset instruction database successfully, the system considers that the user has issued a legal charging instruction, and thus starts the charging authentication process.

[0119] Or, furthermore, when the extracted keyword does not match the instruction in the preset instruction database, the system considers that the user has not issued a legal charging instruction, and thus does not start the charging authentication process and may give corresponding prompts.

[0120] It can be seen that in this embodiment, voice instructions can be effectively recognized and processed, ensuring the safety and accuracy of the charging authentication process, and improving the response ability and user experience of the intelligent system.

[0121] In one embodiment, if so, security authentication is performed on the user to obtain the authentication result corresponding to the user, including: obtaining the voiceprint information in the target voice information; obtaining the multi-dimensional authentication information of the user; performing security authentication on the voiceprint information and the multi-dimensional authentication information to obtain the authentication result corresponding to the user. [[ID=X]]

[0122] Among them, voiceprint information refers to the characteristic information extracted from the voice signal that can represent the identity of the speaker. Voiceprint information is usually related to the physiological structure of the human body and has uniqueness and stability. Voiceprint information includes characteristics such as the frequency, pitch, and speaking speed of the voice.

[0123] Among them, multi-dimensional authentication information refers to various different types of information used for identity authentication, and these information can verify the user's identity from multiple dimensions, improving the security of authentication.

[0124] Specifically, the multi-dimensional authentication information can be voiceprint recognition to analyze the user's voice characteristics, device fingerprint detection of the uniqueness of the device, behavior characteristic analysis of the user's operation habits, and dynamic password (OTP) generation of a one-time password to ensure the security of identity authentication.

[0125] Specifically, the device fingerprint is the unique identification information of the device, which can be generated by collecting the hardware information, system information, network environment information, etc. of the device and is used to identify and distinguish different devices. The behavior characteristics are the operation habits and behavior patterns of the user during the use of the device or application, such as click frequency, sliding speed, input habits, etc. These characteristics can be used to identify and verify the user's identity. The one-time password (OTP) is a one-time password, usually dynamically generated by the system according to factors such as time, events, or the number of uses, and becomes invalid after each use, which is used to improve the security of identity authentication.

[0126] Among them, in the process of obtaining the voiceprint information in the target voice information, voiceprint recognition technology can be used to analyze the target voice information and extract the voiceprint features that can represent the identity of the speaker.

[0127] Among them, voiceprint recognition is a technology that analyzes the voice signal, extracts voiceprint features, and compares them with the voiceprint templates in the database to identify the identity of the speaker.

[0128] Optionally, in the process of performing security authentication on the voiceprint information and the multi-dimensional authentication information, the extracted voiceprint information and the collected multi-dimensional authentication information are input into the security authentication system for verification. The security authentication system will comprehensively analyze the voiceprint information and the multi-dimensional authentication information according to the preset algorithms and rules to determine whether these information match, so as to obtain the authentication result.

[0129] For example, compare the extracted voiceprint features with the voiceprint templates in the database, at the same time verify whether the device fingerprint is legal, analyze whether the user's behavior characteristics match the historical records, and check whether the one-time password entered by the user is correct. If all the information matches, the authentication result is "passed", and the user is allowed to perform the charging operation; if any one of the information does not match, the authentication result is "not passed", and the user's charging request is rejected.

[0130] It can be seen that in this embodiment, by obtaining the voiceprint information in the target voice information and combining multi-dimensional authentication information such as device fingerprint, behavior characteristics, and one-time password, the user's identity is verified more comprehensively and securely, thus ensuring the security of the charging operation, effectively combining biometric technology and various security authentication means, improving the security and reliability of the system, and preventing illegal access and fraud behavior.

[0131] In one embodiment, performing security authentication on the voiceprint information and the multi-dimensional authentication information to obtain the authentication result corresponding to the user includes: obtaining the preset voiceprint information and preset authentication information corresponding to the user; performing a comparison process on the voiceprint information and the preset voiceprint information to obtain a first comparison result; performing a comparison process on the multi-dimensional authentication information and the preset authentication information to obtain a second comparison result; when both the first comparison result and the second comparison result are authentication passed, determining that the authentication result corresponding to the user is authentication passed; or, when the first comparison result and the second comparison result are not both authentication passed, determining that the authentication result corresponding to the user is authentication failed.

[0132] Among them, the preset voiceprint information is the voiceprint feature of the user pre-stored in the system and is used for comparison during identity authentication.

[0133] Among them, the preset authentication information is other multi-dimensional authentication information preset by the system for the user, such as device fingerprint, behavior characteristics, dynamic password, etc., and is used for comparison with the real-time obtained authentication information.

[0134] Among them, the multi-dimensional authentication information is various real-time obtained authentication information, including but not limited to device fingerprint, behavior characteristics, dynamic password, etc.

[0135] Among them, the first comparison result refers to the result obtained after comparing the voiceprint information with the preset voiceprint information, and can be "passed" or "not passed".

[0136] Among them, the second comparison result refers to the result obtained after comparing the multi-dimensional authentication information with the preset authentication information, and can also be "passed" or "not passed".

[0137] Among them, authentication passed means that when all comparison results show that the user information is consistent with the preset information, the system confirms that the user's identity is legal and allows subsequent operations to be executed.

[0138] Among them, authentication failed means that when any comparison result shows that the user information is inconsistent with the preset information, the system deems the user's identity suspicious and refuses to execute subsequent operations.

[0139] Specifically, in the process of performing a comparison process on the voiceprint information and the preset voiceprint information, a voiceprint recognition algorithm is used to calculate the similarity of the two voiceprint features, and it is judged whether they match according to a preset threshold.

[0140] For example, comparing the voiceprint feature of A speaking currently with the voiceprint template in the database, it is found that the similarity reaches 95%, exceeding the preset threshold of 90%, so the first comparison result is "passed".

[0141] Specifically, during the process of comparing the multi-dimensional authentication information with the preset authentication information, each type of authentication information is verified separately, such as checking whether the device fingerprints are consistent, whether the behavioral characteristics match, and whether the dynamic passwords are correct.

[0142] Furthermore, if both comparison results are "passed", the authentication result is "authentication passed"; if any one of the comparison results is "not passed", the authentication result is "authentication not passed".

[0143] It can be seen that in this embodiment, by separately comparing the voiceprint information and the multi-dimensional authentication information and comprehensively judging the comparison results, the user identity can be verified more accurately, ensuring that only legitimate users can access the service, thereby improving the accuracy and security of identity authentication.

[0144] It should be noted that in the above various embodiments, there is not necessarily a certain sequence among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present application that in different embodiments, the above steps can have different execution sequences, that is, they can be executed in parallel or exchanged, etc.

[0145] As another aspect of the embodiments of the present application, the embodiments of the present application provide a charging control device based on voice interaction.

[0146] See Figure 4 , Figure 4 which is a schematic structural diagram of a charging control device based on voice interaction provided by the embodiments of the present application. Applied to a charging control system, as Figure 4 shown, the charging control device 400 based on voice interaction includes:

[0147] An acquisition unit 401, configured to acquire the first voice data of the user;

[0148] A processing unit 402, configured to perform fusion analysis processing on the first voice data to obtain target voice information;

[0149] An authentication unit 403, configured to determine whether to perform charging authentication according to the target voice information;

[0150] The authentication unit 403 is further configured to, if so, perform security authentication on the user to obtain the authentication result corresponding to the user;

[0151] A control unit 404, configured to perform charging control according to the authentication result and the target voice information.

[0152] In this embodiment, by obtaining the user's first voice data and performing fusion analysis, the user can complete the charging operation through voice commands, reducing the dependence on the mobile phone APP or the charging pile screen, realizing effective voice interaction and coordination among the charging pile, the mobile phone APP and the smart speaker, ensuring timely data synchronization among devices, avoiding the problem of information silos, simplifying the operation process. The voice interaction method enables the user to charge without manual operation when driving or carrying items. Especially in situations where manual operation is inconvenient, such as driving, carrying items, bad weather or at night, voice control provides a more convenient and safe way of use. Further, before charging authentication, user security authentication is performed through the target voice information to ensure that only authenticated users can perform the charging operation, improving the security of the system. Through the fusion analysis of voice data, the interaction method can be adaptively adjusted in different scenarios, enhancing the coordination among the charging pile, the mobile phone APP and the smart speaker, and providing a more consistent and seamless user experience. Therefore, through the automated voice recognition and authentication process in this embodiment, errors caused by the complexity of manual operation can be reduced, and the success rate and efficiency of the charging operation are improved.

[0153] In one embodiment, when performing fusion analysis and processing on the first voice data to obtain the target voice information, the processing unit 402 is further configured to: perform voice recognition on the first voice data to obtain first voice information; obtain at least one set of fusion parameters; and perform fusion analysis and processing on the first voice information according to the at least one set of fusion parameters to obtain the target voice information.

[0154] In one embodiment, when performing voice recognition on the first voice data to obtain first voice information, the processing unit 402 is further configured to: perform data preprocessing on the first voice data to obtain the preprocessed first voice data; extract feature parameters from the preprocessed first voice data to obtain first feature parameters; input the first feature parameters into a preset acoustic model for processing to obtain second voice data output by the preset acoustic model; and input the second voice data into a preset language model for processing to obtain the first voice information output by the preset language model.

[0155] In one embodiment, when performing fusion analysis processing on the first voice information according to the at least one set of fusion parameters to obtain target voice information, the processing unit 402 is further configured to: analyze the at least one set of fusion parameters to obtain the parameter type of each set of fusion parameters; determine a fusion mode according to the parameter type; perform fusion processing on the at least one set of fusion parameters according to the fusion mode to obtain first fusion data; extract features from the first fusion data to obtain second feature parameters; input the second feature parameters into a preset classification model for scene recognition to obtain a first scene label output by the preset classification model; and perform analysis processing on the first voice information according to the first scene label to obtain target voice information.

[0156] In one embodiment, when performing analysis processing on the first voice information according to the first scene label to obtain target voice information, the processing unit 402 is further configured to: perform scene recognition on the first voice information to obtain a second scene label; determine whether the first scene label and the second scene label are within the same label range; if so, determine that the first voice information is target voice information; or, if not, perform prediction processing on the first voice information according to the first scene label to obtain target voice information.

[0157] In one embodiment, when determining whether to perform charging authentication according to the target voice information, the authentication unit 403 is further configured to: extract multiple keywords from the target voice information; determine whether the multiple keywords exist in a preset instruction database; if so, determine to perform charging authentication; or, if not, determine not to perform charging authentication.

[0158] In one embodiment, when, if so, performing security authentication on the user to obtain an authentication result corresponding to the user, the authentication unit 403 is further configured to: obtain voiceprint information in the target voice information; obtain multi-dimensional authentication information of the user; and perform security authentication on the voiceprint information and the multi-dimensional authentication information to obtain an authentication result corresponding to the user.

[0159] In one embodiment, in the process of performing security authentication on the voiceprint information and the multi-dimensional authentication information to obtain the authentication result corresponding to the user, the authentication unit 403 is further configured to: obtain the preset voiceprint information and preset authentication information corresponding to the user; perform a comparison process on the voiceprint information and the preset voiceprint information to obtain a first comparison result; perform a comparison process on the multi-dimensional authentication information and the preset authentication information to obtain a second comparison result; when both the first comparison result and the second comparison result are authentication passed, determine that the authentication result corresponding to the user is authentication passed; or, when the first comparison result and the second comparison result are not both authentication passed, determine that the authentication result corresponding to the user is authentication failed.

[0160] It should be noted that the above-mentioned charging control device based on voice interaction can execute the charging control method based on voice interaction provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in the embodiments of the charging control device based on voice interaction, reference can be made to the charging control method based on voice interaction provided by the embodiments S10-S50 of the present application.

[0161] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the charging control method based on voice interaction as described in the foregoing embodiments.

[0162] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0163] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A charging control method based on voice interaction, characterized in that, Applied to a charging control system, the method includes: Obtain the first voice data of the user; Perform fusion analysis processing on the first voice data to obtain target voice information; Judge whether to perform charging authentication according to the target voice information; If so, perform security authentication on the user to obtain the authentication result corresponding to the user; Perform charging control according to the authentication result and the target voice information.

2. The method according to claim 1, characterized in that, The performing fusion analysis processing on the first voice data to obtain target voice information includes: Perform voice recognition on the first voice data to obtain first voice information; Obtain at least one set of fusion parameters; According to the at least one set of fusion parameters, perform fusion analysis processing on the first voice information to obtain target voice information.

3. The method according to claim 2, characterized in that, The performing voice recognition on the first voice data to obtain first voice information includes: Perform data preprocessing on the first voice data to obtain the preprocessed first voice data; Extract feature parameters from the preprocessed first voice data to obtain first feature parameters; Input the first feature parameters into a preset acoustic model for processing to obtain second voice data output by the preset acoustic model; Input the second voice data into a preset language model for processing to obtain first voice information output by the preset language model.

4. The method according to claim 2, wherein The performing, according to the at least one set of fusion parameters, fusion analysis processing on the first voice information to obtain target voice information includes: Analyze the at least one set of fusion parameters to obtain the parameter type of each set of fusion parameters; Determine the fusion mode according to the parameter type; Perform fusion processing on the at least one set of fusion parameters according to the fusion mode to obtain first fusion data; Extract feature parameters from the first fusion data to obtain second feature parameters; Input the second feature parameters into a preset classification model for scene recognition to obtain a first scene label output by the preset classification model; Analyze and process the first voice information according to the first scene label to obtain target voice information.

5. The method according to claim 4, characterized in that The analyzing and processing the first voice information according to the first scene label to obtain target voice information includes: Perform scene recognition on the first voice information to obtain a second scene label; Judge whether the first scene label and the second scene label are within the same label range; If so, determine that the first voice information is target voice information; or, If not, perform prediction processing on the first voice information according to the first scene label to obtain target voice information.

6. The method according to claim 1, characterized in that, The judging whether to perform charging authentication according to the target voice information includes: Extract multiple keywords from the target voice information; Judge whether the multiple keywords exist in a preset instruction database; If they exist, determine to perform charging authentication; or, If they do not exist, determine not to perform charging authentication.

7. The method according to claim 1, wherein The if so, performing security authentication on the user to obtain the authentication result corresponding to the user includes: Obtain the voiceprint information in the target voice information; Obtain the multi-dimensional authentication information of the user; Perform security authentication on the voiceprint information and the multi-dimensional authentication information to obtain the authentication result corresponding to the user.

8. The method according to claim 7, wherein The performing security authentication on the voiceprint information and the multi-dimensional authentication information to obtain the authentication result corresponding to the user includes: Obtain the preset voiceprint information and preset authentication information corresponding to the user; Perform a comparison process on the voiceprint information and the preset voiceprint information to obtain a first comparison result; Perform a comparison process on the multi-dimensional authentication information and the preset authentication information to obtain a second comparison result; When both the first comparison result and the second comparison result are authentication passed, determine that the authentication result corresponding to the user is authentication passed; or, When the first comparison result and the second comparison result are not both authentication passed, determine that the authentication result corresponding to the user is authentication failed.

9. A charging control system, comprising: A memory, a processor, and a computer program stored on the memory and running on the processor, where when the computer program is executed by the processor, it implements the voice interaction-based charging control method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the voice interaction-based charging control method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Charging pile identity authentication method and device and electronic equipment

    CN109493870A

  • Charging method and system based on voiceprint recognition

    CN109544803A

  • Wireless starting method and system of intelligent charging pile

    CN116564005A