A public place intelligent interactive robot control method and an interactive robot
By working in tandem with the central control console and the robot system, user needs are identified in real time and the optimal robot is matched, which solves the efficiency and experience problems of multi-user navigation services in large public places, realizes personalized and fast navigation services, and improves user satisfaction.
Patent Information
- Application Number
- CN202411900660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-23
AI Technical Summary
In large public places, robots cannot meet the time-limited navigation needs of multiple users at the same time, resulting in a poor user experience. In particular, they cannot provide navigation services in a timely manner when time is tight, which affects subsequent travel arrangements.
Through the collaborative work of the central control console and the robot, the system acquires facial images and voice information from the user terminal in real time, performs semantic and facial recognition, determines the user's level and itinerary information, matches the optimal robot, generates path control schemes, and provides personalized and fast navigation services.
This effectively avoids excessively long waiting times for users, improves service efficiency and security, enhances user interaction convenience, provides personalized services, ensures that the subsequent travel plans of users in urgent situations are not affected, and improves user experience and satisfaction.
Smart Images

Figure CN119759010B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a control method and interactive robot for intelligent interactive robots in public places. Background Technology
[0002] With the rapid development of science and technology and the increasing maturity of artificial intelligence, robots, as an important carrier of this technology, are gradually integrating into people's daily lives. Especially in public places, the application of robots makes public places more intelligent, convenient and efficient, bringing a brand-new service experience to the public.
[0003] Currently, robots in public places such as airports and shopping malls mainly provide navigation and guidance services, helping users quickly find their target locations based on their needs and offering guided tours. They obtain the user's destination and guide them along a generated path. However, in these public places, especially when the area is large and there is high foot traffic, multiple users often require guidance services. In such cases, because the robot's built-in program can only serve one user's need, it cannot meet all user requirements, resulting in a poor user experience. This is particularly true in public places like airports and train stations, where users need time-limited robot navigation services. For example, if a user needs to reach a location such as a boarding gate or station ticket gate within a very limited time due to an approaching departure time for a flight, train, or other transportation, the required robot navigation service must be completed within a specific time constraint. However, because the robot cannot provide timely navigation services to other users, it may affect the user's subsequent travel arrangements, severely reducing the user's overall experience with robots in public places.
[0004] Therefore, a solution to the above problems is urgently needed. Summary of the Invention
[0005] This application provides a control method and interactive robot for intelligent interactive robots in public places, which can improve the user experience of using the robot.
[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0007] Firstly, a method for controlling intelligent interactive robots in public places is provided, applied to a central control console, wherein the central control console is connected to at least one robot, and each robot is equipped with a first camera. The method includes:
[0008] In response to receiving path navigation instructions from multiple user terminals, the user terminal corresponding to the earliest path navigation instruction is taken as the target terminal.
[0009] In response to receiving voice information sent by the target terminal, a shooting command is sent to the target terminal. The shooting command is used to activate the second camera of the target terminal so that the target terminal can capture a facial image of the target user corresponding to the target terminal.
[0010] In response to receiving the face image sent by the target terminal, semantic recognition is performed on the voice information to determine the first target location that the target user needs to reach, and face recognition is performed on the face image to obtain the travel information of the target user;
[0011] The robot information of each robot is acquired in real time, and the first current location of the target user is acquired in real time.
[0012] Based on the first target location and the travel information, the user level of the target user is determined;
[0013] Based on the user level, the first target location, each robot information and the first current location, determine the optimal matching robot and obtain the second current location of the optimal matching robot;
[0014] Based on the first target location, the first current location, and the second current location, the optimal matching robot path control scheme is generated and displayed.
[0015] In one possible implementation of the first aspect, generating the path control scheme for the optimal matching robot based on the first target location, the first current location, and the second current location includes:
[0016] A first path is generated based on the second current location, the first current location, and the first target location. The first path includes the second current location, the first current location, and the first target location. The second current location is the first waypoint of the first path, the first current location is the second waypoint of the first path, and the first target location is the third waypoint of the first path.
[0017] Based on the first path, a first control command is generated and sent to the optimal matching robot so that the optimal matching robot executes the first control command.
[0018] In one possible implementation of the first aspect, the method further includes:
[0019] A first rendezvous point is generated based on the second current location and the first target location;
[0020] A second path is generated using the first current location and the first meeting point. The second path includes the first current location and the first meeting point. The first current location is the first waypoint of the second path, and the first meeting point is the second waypoint of the second path.
[0021] Send the second path to the target terminal;
[0022] A third path is generated using the first rendezvous point and the second current location. The third path includes the first rendezvous point and the second current location. The second current location is the first waypoint of the third path, and the first rendezvous point is the second waypoint of the third path.
[0023] Based on the third path, a second control command is generated and sent to the optimal matching robot so that the optimal matching robot executes the second control command;
[0024] A fourth path is generated using the first meeting point and the first target location. The fourth path includes the first meeting point and the first target location. The first meeting point is the first waypoint of the fourth path, and the first target location is the second waypoint of the fourth path.
[0025] Based on the fourth path, a third control command is generated and sent to the optimal matching robot so that the optimal matching robot executes the third control command.
[0026] In one possible implementation of the first aspect, prior to generating the third control instruction according to the fourth path, the method includes:
[0027] The arrival time of the target user is determined using the second path;
[0028] The preset waiting time threshold for the optimal matching robot is set based on the arrival time;
[0029] Within a preset waiting time threshold, the first camera of the optimal matching robot is controlled to capture images;
[0030] The similarity between the image and the face image is calculated. If the similarity is greater than a preset similarity threshold, the optimal matching robot executes a third control command.
[0031] If the waiting time of the optimal matching robot is greater than a preset waiting time threshold and the similarity is less than a preset similarity threshold, then the optimal matching robot sends a signal to the central control console to exit the current service.
[0032] In response to receiving the signal to exit the current service, a target parking area closest to the optimal matching robot is determined in the preset parking area, and the location of the parking area is sent to the optimal matching robot so that the optimal matching robot can go to the parking area location.
[0033] In one possible implementation of the first aspect, the step of performing semantic recognition on the voice information to determine the first target location that the target user needs to reach includes:
[0034] The voice information is converted into a voice signal;
[0035] Endpoint detection is performed on the speech signal to determine the start and end endpoints of the speech signal;
[0036] An audio segment is obtained through the start endpoint and the end endpoint;
[0037] The audio segment is processed using a preset filtering algorithm to obtain a filtered audio segment;
[0038] The filtered audio segment is input into a pre-trained acoustic model to obtain the feature vector of the audio segment;
[0039] The filtered audio segment is input into a pre-trained language model to obtain a word sequence of the audio segment;
[0040] By using the feature vector and the word sequence, semantic understanding is performed on the audio segment to obtain the key information of the audio segment;
[0041] Based on the key information, the first target location that the target user needs to reach is determined.
[0042] In one possible implementation of the first aspect, the endpoint detection of the speech signal to determine the start and end endpoints of the speech signal includes:
[0043] The speech signal is segmented into frames to obtain multiple speech frames of the speech signal;
[0044] Perform a Fourier transform on each of the speech frames to obtain the spectrogram of each speech frame;
[0045] Calculate the spectral energy value for each of the aforementioned spectrograms;
[0046] When the spectral energy value is less than a preset low threshold and the spectral energy value of the speech frame adjacent to the speech frame whose spectral energy value is less than the preset low threshold is greater than a preset high threshold, the first speech frame whose spectral energy value is less than the preset low threshold is taken as the start endpoint of the speech signal.
[0047] When the spectral energy value is less than a preset high threshold and the spectral energy value of the speech frame adjacent to the speech frame whose spectral energy value is less than the preset high threshold is greater than the preset high threshold, the last speech frame whose spectral energy value is greater than the preset high threshold is taken as the end point of the speech signal.
[0048] In one possible implementation of the first aspect, determining the user level of the target user based on the first target location and the travel information includes:
[0049] Based on the travel information, the first time point of the target user is determined;
[0050] Based on the first target location and the first current location, determine the path length from the first current location to the first target location;
[0051] Get the current time point;
[0052] Based on the path length and the preset speed of the optimal matching robot, calculate the time period for the optimal matching robot to reach the first target location, and based on the current time node and the time period, calculate the second time node for the optimal matching robot to reach the first target location.
[0053] The first path is mapped into a preset geographic data model, the influencing factors of the first path are determined, and the third time node for the optimal matching robot to reach the first target location is determined based on the influencing factors.
[0054] Using the grey relational analysis algorithm, an evaluation matrix is constructed from the first time node, the second time node, and the third time node;
[0055] Determine the weight values of the first time node, the second time node, and the second time node;
[0056] The weight values and the evaluation matrix are combined using a preset synthesis operator to obtain a comprehensive evaluation result of the user level.
[0057] The user level of the target user is determined based on the comprehensive evaluation results of the user level.
[0058] In one possible implementation of the first aspect, the robot information includes robot device information, and determining the optimal matching robot based on the user level, the first target location, and the robot information includes:
[0059] Obtain spatial location data of the robot;
[0060] The target user, the first target location, and the robot are mapped into the geographic data model and displayed.
[0061] Based on the user level, a time threshold for the user is determined in a preset user level database. The time threshold includes a first time period and a second time period. The first time period is the time period between the first current location and the first target location, and the second time period is the time period between the second current location and the first current location.
[0062] The target robot is retrieved from the geographic data model using the time threshold.
[0063] The environmental information of the public place is determined using the geographic data model.
[0064] By using the environmental information of the public place, the time required for the target robot to reach the first target location by passing through the location of the target user and the pre-consumed power can be determined in real time.
[0065] The electrical energy reserve of the target robot is determined by the robot equipment information of the target robot.
[0066] When the target robot's energy reserves are greater than the pre-consumed energy and the time required for the target robot to reach the first target location via the first current location is less than a preset time threshold, the optimal matching robot is determined.
[0067] In one possible implementation of the first aspect, after determining the first target location that the target user needs to reach, and performing facial recognition on the facial image to obtain the target user's travel information, the method further includes:
[0068] Retrieve the second target location closest to the target user within the geographic data model;
[0069] Based on the target user's travel information, the target user's destination is obtained, and it is determined whether the first target location and the second target location are the destination;
[0070] If both the first target location and the second target location are the destination, then a second target location confirmation message is generated and sent to the target terminal based on the second target location.
[0071] After receiving the confirmation information sent by the target terminal, the second target location is taken as the first target location;
[0072] Upon receiving the negative information sent by the target terminal, the steps of acquiring robot information of each robot in real time and acquiring the first current location of the target user in real time are executed.
[0073] Secondly, this application provides an interactive robot for use with robots, each of which is connected to a central control console, and each of which is equipped with a first camera. The interactive robot also includes:
[0074] The memory is configured to store instructions; and
[0075] The processor is configured to retrieve and execute the instructions from the memory.
[0076] By receiving route navigation instructions from multiple user terminals and designating the earliest received instruction as the target terminal, the problem of excessively long user wait times can be effectively avoided, improving service efficiency. Identity verification via facial recognition not only enhances security but also increases the convenience of user interaction with the robot, simplifying the process and improving the user experience. After receiving the facial image from the target terminal, semantic recognition of the voice information determines the first destination the target user needs to reach. Semantic recognition technology accurately identifies user instructions, better understands user needs and intentions, and provides personalized services, enhancing the user experience. By acquiring the user's travel information and determining the user's urgency level based on their profile, more comprehensive services can be provided to the target user, ensuring that their subsequent travel arrangements are not affected. Matching the target user with the optimal robot provides faster route planning. Based on the user's and robot's locations, a route adjustment scheme is determined, providing clear and accurate navigation guidance and improving user satisfaction.
[0077] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description
[0078] Figure 1 A flowchart illustrating a control method for an intelligent interactive robot in a public place, provided as an embodiment of this application;
[0079] Figure 2 A schematic diagram illustrating a target location determination process provided in an embodiment of this application;
[0080] Figure 3 A flowchart illustrating another intelligent interactive robot control method for public places provided in this application embodiment;
[0081] Figure 4This is a schematic diagram of a first assembly point confirmation provided in an embodiment of this application. Detailed Implementation
[0082] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0083] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0084] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0085] Figure 1 The illustration shows a flowchart of a control method for an intelligent interactive robot in a public place according to an embodiment of this application. Figure 1 As shown in the figure, this application provides a method for controlling an intelligent interactive robot in a public place, which may include the following steps.
[0086] S110. In response to receiving path navigation instructions from multiple user terminals, the user terminal corresponding to the earliest path navigation instruction is taken as the target terminal.
[0087] S120. In response to receiving voice information sent by the target terminal, a shooting command is sent to the target terminal. The shooting command is used to call the second camera of the target terminal so that the target terminal can capture a face image of the target user corresponding to the target terminal.
[0088] S130. In response to receiving a face image sent by the target terminal, perform semantic recognition on the voice information to determine the first target location that the target user needs to reach, and perform face recognition on the face image to obtain the target user's travel information.
[0089] S140. Obtain robot information for each robot in real time, and obtain the first current location of the target user in real time;
[0090] S150. Determine the user level of the target user based on the first target location and travel information;
[0091] S160. Based on the user level, the first target location, each robot's information, and the first current location, determine the optimal matching robot and obtain the second current location of the optimal matching robot;
[0092] S170. Based on the first target location, the first current location, and the second current location, generate and display the optimal path control scheme for the matching robot.
[0093] Figure 3 This illustration schematically shows a flowchart of another intelligent interactive robot control method for public places according to an embodiment of this application. Figure 3 As shown in the flowchart, the control method for intelligent interactive robots in public places includes a target terminal 10, a central control station 20, and a robot 30. The central control station sends verification information to the target terminal. After the target user agrees to undergo facial recognition verification, the central control station obtains the target user's travel information and the first target location, determines the user level, and sends the user level to the robot.
[0094] In response to receiving path navigation instructions from multiple user terminals, the user terminal corresponding to the earliest path navigation instruction is designated as the target terminal. In this embodiment, the path navigation instruction refers to sending a target station to the robot, which then guides the user to the target station according to a pre-set route. The target terminal refers to various devices used by the user; in this embodiment, the target terminal can be a mobile phone, tablet computer, or other similar device. In other words, when path navigation instructions are received from multiple user terminals, such as mobile phones, the mobile phone that first issued the path navigation instruction is designated as the target terminal.
[0095] After identifying the target terminal, when the voice information sent by the target terminal is received, a shooting command is sent to the target terminal. The target user using the target terminal is photographed through the second camera of the target terminal to obtain a facial image of the target user corresponding to the target terminal. In this embodiment, the shooting command is used to control the second camera of the target terminal to photograph the facial image of the target user.
[0096] Upon receiving the user's facial image from the target terminal, semantic recognition of the user's voice information is performed to determine the first target location the user needs to reach. In this embodiment, the first target location is determined by recognizing the user's voice information. Semantic recognition of the user's voice information can be achieved using a Hidden Markov Model (HMM). An HMM is a temporal probabilistic model that can be used to model speech signals and identify words or phonemes in speech by learning probability distributions. Using an HMM, keywords in the user's voice information can be identified, and these keywords determine the first target location the user needs to reach. Upon receiving the user's facial image, facial recognition is performed. Facial recognition is a technology for identity verification based on human facial features. It utilizes computer vision and pattern recognition technologies to detect, extract, and compare faces in facial images or videos, thereby achieving automatic verification and identification of an individual's identity. Once the user agrees to facial recognition and passes the verification, the target user's travel information is obtained.
[0097] After obtaining the target user's travel information, the robot information for each robot is acquired in real time. This can be achieved by equipping the robots with various sensors, such as vision sensors, force sensors, sound sensors, and temperature sensors. Vision sensors can be used to monitor the robot's surrounding environment in real time, as well as to identify and track objects; force sensors can measure the forces between the robot and other objects, such as humans, other robots, or objects in the work environment; sound sensors can monitor ambient sounds, helping to understand the dynamic situation around the robot; and temperature sensors can monitor ambient temperature, helping the robot adapt to different working environments. Simultaneously with acquiring the robot information for each robot in real time, the target user's initial current location is also obtained. This can be determined by identifying the location of the target user's target terminal. The location of the target user's target terminal can be determined using GPS positioning, a global radio satellite navigation system. For example, if the target user's target terminal is a mobile phone, its built-in GPS module can be used to obtain precise location information from the phone, thus determining the target user's initial current location.
[0098] Based on the first target location and travel information, the user level of the target user is determined. In this embodiment, the user level is determined by analyzing the time urgency of the target user's journey from the first current location to the first target location. Travel information refers to detailed information related to the target user's transportation. That is, the first time node of the target user is determined through travel information. In this embodiment, the first time node refers to the latest boarding time of the target user's mode of transportation. For example, if the target user's mode of transportation is a train, the first time node is the latest time node to board the train. After determining the first time node of the target user, the time between the first target location and the first current location is determined and this time is taken as the second time node. Subsequently, the time between the arrival of the optimal matching robot at the first target location and the time delay caused by environmental factors during this process are taken as the third time node. An evaluation matrix is constructed using the grey relational analysis algorithm, which is a statistical analysis method based on grey system theory and is mainly used to evaluate the correlation between multiple systems or variables. The weight values for the first, second, and third time nodes are determined using the Analytic Hierarchy Process (AHP). AHP is a decision analysis tool that constructs a hierarchical model, performs pairwise comparisons, and calculates weight values. Next, a preset synthesis operator is used to synthesize the weight values with the evaluation matrix to obtain a comprehensive user level evaluation result. In this embodiment, the preset synthesis operator can be a weighted average method, a mathematical method used to calculate the average value of a set of data while considering the importance or weight of each data point. Based on the comprehensive user level evaluation result, the user level of the target user is determined.
[0099] After determining the target user's user level, the optimal matching robot is determined based on the user level, the first target location, each robot's information, and the first current location. In other words, determining the optimal matching robot requires comprehensively considering multiple factors, including the user level, the first target location, and specific information about each robot, such as its current power reserves. Specifically, the latest departure time of the target user is determined based on the user level. Then, the time required for the robot to travel from the target user's location to the first target location, as well as the power consumption, are determined. A search is performed at the target user's location, and robots with power reserves greater than the required power consumption, and whose travel time from the target user's location to the first target location is less than the target user's latest departure time, are selected as the optimal matching robots. After determining the optimal matching robot, its second current location is obtained. In this embodiment, the second current location is the location information of the optimal matching robot. Using wireless LAN technology, the received WiFi signal strength or phase information is measured and compared with WiFi access points at known locations to determine the robot's location.
[0100] Based on the first target location, the first current location, and the second current location, an optimal path control plan for the matching robot is generated and displayed. In other words, the optimal matching robot's path control plan is generated and displayed on the target user's target terminal using the target user's location, the optimal matching robot's location, and the first target location. Specifically, different path control plans can be generated based on the target user's location, the optimal matching robot's location, and the first target location. For example, when the optimal matching robot is located between the target user and the first target location, a first meeting point can be generated. After the target user and the optimal matching robot reach the first meeting point, they will proceed together to the first target location. Alternatively, when the target user is located between the optimal matching robot and the first target location, the optimal matching robot can reach the target user's location before proceeding together to the first target location. After generating the optimal matching robot's path control plan, displaying it on the target user's target terminal can be achieved through a mobile application. For example, a mobile application can be developed that integrates the path control plan. Users can download and install the application to view the robot's real-time location and path planning on their mobile devices. The mobile application can support real-time updates and push notifications so that users can obtain the latest path control information promptly.
[0101] By receiving route navigation instructions from multiple user terminals and designating the earliest received instruction as the target terminal, the problem of excessively long user wait times can be effectively avoided, improving service efficiency. Identity verification via facial recognition not only enhances security but also increases the convenience of user interaction with the robot, simplifying the process and improving the user experience. After receiving the facial image from the target terminal, semantic recognition of the voice information determines the first destination the target user needs to reach. Semantic recognition technology accurately identifies user instructions, better understands user needs and intentions, and provides personalized services, enhancing the user experience. By acquiring the user's travel information and determining the user's urgency level based on their profile, more comprehensive services can be provided to the target user, ensuring that their subsequent travel arrangements are not affected. Matching the target user with the optimal robot provides faster route planning. Based on the user's and robot's locations, a route adjustment scheme is determined, providing clear and accurate navigation guidance and improving user satisfaction.
[0102] In one embodiment of this example, based on the first target location, the first current location, and the second current location, a path control scheme for the optimal matching robot is generated, including:
[0103] S210. Based on the second current location, the first current location, and the first target location, generate a first path. The first path includes the second current location, the first current location, and the first target location. The second current location is the first waypoint of the first path, the first current location is the second waypoint of the first path, and the first target location is the third waypoint of the first path.
[0104] S220. Based on the first path, generate a first control command and send the first control command to the optimal matching robot so that the optimal matching robot executes the first control command.
[0105] Based on the first target location, the first current location, and the second current location, a path control scheme for the optimal matching robot is generated. In this embodiment, the first target location is the location the target user needs to reach, the first current location is the target user's current location, and the second current location is the location of the optimal matching robot. Specifically, firstly, a first path is generated based on the second current location, the first current location, and the first target location. In this embodiment, the first path is the path from the location of the optimal matching robot through the location of the target user to the first target location. The second current location is the first waypoint of the first path, the first current location is the second waypoint of the first path, and the first target location is the third waypoint of the first path. That is, when the first current location is between the first target location and the second current location, the distance from the location of the optimal matching robot to the first target location is greater than the distance from the target user to the first target location. Therefore, the target user needs to wait for the optimal matching robot to reach their current location before proceeding together to the first target location.
[0106] After obtaining the first path, a first control command is generated based on the first path. In this embodiment, the first control command is to make the optimal matching robot move along the first path. That is, after the first path is generated, the first control command is issued to the optimal matching robot so that the optimal matching robot moves along the first path, first from the location of the optimal matching robot to the location of the target user, and finally to the first target location.
[0107] By generating a first path using the second current location, the first current location, and the first target location, and sending the first control command to the optimally matched robot, the robot's movement time and energy consumption can be reduced, significantly improving the robot's working efficiency and accuracy, and enhancing the user's interaction experience with the robot.
[0108] In one embodiment of this invention, the method further includes:
[0109] S310. Generate a first rendezvous point based on the second current location and the first target location;
[0110] S320. Generate a second path using the first current location and the first meeting point. The second path includes the first current location and the first meeting point. The first current location is the first waypoint of the second path, and the first meeting point is the second waypoint of the second path.
[0111] S330, Send the second path to the target terminal;
[0112] S340. Generate a third path using the first meeting point and the second current location. The third path includes the first meeting point and the second current location. The second current location is the first waypoint of the third path, and the first meeting point is the second waypoint of the third path.
[0113] S350. Based on the third path, generate a second control command and send the second control command to the optimal matching robot so that the optimal matching robot executes the second control command;
[0114] S360. Generate a fourth path using the first meeting point and the first target location. The fourth path includes the first meeting point and the first target location. The first meeting point is the first waypoint of the fourth path, and the first target location is the second waypoint of the fourth path.
[0115] S370. Based on the fourth path, generate a third control command and send the third control command to the optimal matching robot so that the optimal matching robot executes the third control command.
[0116] Based on the second current location and the first target location, a first meeting point is generated. In this embodiment, the first meeting point is the location where the target user and the optimal matching robot gather or concentrate. That is, a meeting point is generated between the location of the optimal matching robot and the first target location, where the target user and the optimal matching robot meet. When the first target location is between the second current location and the location of the target user, the target user and the optimal matching robot simultaneously head to the first meeting point and then reach the first target location together, which can effectively save the target user's time. Figure 4 As shown in the diagram, this application provides a structural schematic diagram for determining a first assembly point. Figure 4 As shown, point A is the location of the optimal matching robot, point B is the location of the target user, point C is the first target location, and point D is the first meeting point. When the target user's location is in the middle between the location of the optimal matching robot and the target user's location, the first meeting point is generated in the middle between the target user's location and the optimal matching robot's location, reducing the time it takes for the target user to reach the first target location.
[0117] A second path is generated using the first current location and the first meeting point. This can be achieved using Dijkstra's algorithm, an algorithm for finding the single-source shortest path in a weighted graph. In this embodiment, the second path is the path from the target user to the first meeting point. The first current location is the first waypoint of the second path, and the first meeting point is the second waypoint of the second path. In other words, after determining the first meeting point, the second path is generated using the target user's location and the first meeting point, and then sent to the target user's target terminal. The second path guides the user to the first meeting point.
[0118] Sending the second path to the target terminal can be done via network transmission. Network transmission refers to the process of data being transmitted from the sender to the receiver in a computer network. The second path can be transmitted via network protocols. The sender encapsulates the second path into a data packet and sends it to the target terminal via the network. After receiving the data packet, the target terminal decapsulates it to obtain the second path and sends it to the target terminal. The target user can then reach the first meeting point based on the second path in the target terminal.
[0119] A third path is generated using the first meeting point and the second current location. In this embodiment, the third path is the path from the second current location where the optimal matching robot is located to the first meeting point. The second current location is the first waypoint of the third path, and the first meeting point is the second waypoint of the third path. In other words, the path from the optimal matching robot to the first meeting point is determined by the second current location where the optimal matching robot is currently located and the first meeting point, so that the optimal matching robot and the target user meet at the first meeting point and arrive at the first target location together.
[0120] Based on the third path, a second control command is generated. In this embodiment, the second control command is used to make the optimal matching robot move according to the third path. That is, after the third path is determined, the second control command is issued to the optimal matching robot so that the optimal matching robot reaches the first meeting point and then guides the target user to the first target location.
[0121] After issuing the second control command to the optimal matching robot, a fourth path is generated through the first meeting point and the first target location. In this embodiment, the fourth path refers to the path from the first meeting point to the first target location. The first meeting point is the first waypoint of the fourth path, and the first target location is the second waypoint of the fourth path. That is to say, after the optimal matching robot and the target user meet at the first meeting point, the optimal matching robot guides the target user to the first target location.
[0122] After determining the fourth path, a third control command is generated based on the fourth path. In this embodiment, the third control command is used to make the optimal matching robot move according to the fourth path. That is, after the fourth path is determined, the third control command is issued to the optimal matching robot so that after the optimal matching robot reaches the first meeting point, it guides the target user to the first target location.
[0123] By using the first target location, the first current location, and the second current location, a path control scheme for the optimal matching robot is generated. This enables the optimal matching robot to reach the destination more quickly and accurately, thereby improving the overall task execution efficiency and enhancing the user experience of the robot.
[0124] In one embodiment of this example, before generating the third control command according to the fourth path, the following steps are included:
[0125] S410. Determine the arrival time of the target user through the second path;
[0126] S420. Set the preset waiting time threshold for the optimal matching robot based on the arrival time;
[0127] S430. Within the preset waiting time threshold, control the first camera of the optimal matching robot to capture images;
[0128] S440. Calculate the similarity between the image and the face image. If the similarity is greater than the preset similarity threshold, the optimal matching robot executes the third control instruction.
[0129] S450. If the waiting time of the optimal matching robot is greater than the preset waiting time threshold and the similarity is less than the preset similarity threshold, the optimal matching robot sends a signal to the central control station to exit the current service.
[0130] S460, In response to receiving a signal to exit the current service, determine the target parking area closest to the optimal matching robot in the preset parking area, and send the parking area location to the optimal matching robot so that the optimal matching robot can go to the parking area location.
[0131] Before generating the third control command based on the fourth path, it is necessary to determine whether the target user has arrived at the first meeting point. When the target user arrives at the first meeting point, the optimal matching robot executes the third control command to bring the target user to the first target location. If the target user has not arrived at the first meeting point, the optimal matching robot sends a signal to exit the current service to the central control station, and then the optimal matching robot returns to the parking area. Specifically, firstly, the arrival time of the target user is determined through the second path. In this embodiment, the second path is the path from the first current location to the first meeting point. That is, the time for the target user to arrive at the first meeting point can be determined by obtaining the path length from the first current location to the first meeting point and the user's average walking speed. The user's average walking speed can be obtained from a preset database. The time for the target user to arrive at the first meeting point can be obtained by dividing the path length by the user's average walking speed.
[0132] After obtaining the arrival time of the target user, a preset waiting time threshold for the optimal matching robot is set based on the arrival time. In this embodiment, the preset waiting time threshold is the time period from the first current location to the first meeting point of the target user. That is, after determining the time period from the first current location to the first meeting point of the target user, this time period is used as the preset waiting time threshold for the optimal matching robot. The optimal matching robot will arrive at the first meeting point in advance and wait for the target user at the first meeting point for the duration of the wait for the target user, which is the preset waiting time threshold.
[0133] While the optimal matching robot is waiting for the target user, the robot's first camera is controlled to capture images. In other words, within a preset waiting time threshold, the optimal matching robot will control its first camera to capture images while waiting for the target user, and store the captured images in the robot's built-in memory. Then, the images stored in the memory are retrieved and compared to determine whether the target user has arrived.
[0134] After the first camera of the optimal matching robot captures an image, a similarity calculation is performed between the image and the face image. This can be done by calculating the Euclidean distance between the two images. Euclidean distance is a straight-line distance between two points in n-dimensional space, or the natural length between two points. It is one of the most common distance metrics, typically used to calculate the similarity or difference between vectors. By calculating the straight-line distance between each pixel in the image and the face image (i.e., the square root of the sum of the squares of the differences between each pixel in the image), the pixel differences between the two images can be compared to determine whether the target user has reached the first meeting point. When the similarity is greater than a preset similarity threshold (in this embodiment, the preset similarity threshold can be determined based on actual conditions), the smaller the pixel distance between the image and the face image, the more similar they are. This means that the optimal matching robot has captured an image of the target user, indicating that the target user has reached the first meeting point. Once it is determined that the target user has reached the first meeting point, the optimal matching robot executes a third control command, meaning that the optimal matching robot will take the target user to the first target location.
[0135] If the waiting time of the optimal matching robot exceeds a preset waiting time threshold, and the similarity is less than a preset similarity threshold, the optimal matching robot sends a signal to the central control station to exit the current service. In other words, if the optimal matching robot waits for the target user at the first meeting point for longer than the preset waiting time threshold, it means the target user did not arrive at the first meeting point on time within the preset waiting time threshold, and the similarity between the captured image and the target user's facial image is less than the preset similarity threshold, meaning the optimal matching robot did not capture or capture the target user's facial image, it indicates that the target user has not yet arrived at the first meeting point. Once it is determined that the target user has not arrived at the first meeting point, the optimal matching robot will send a signal to the central control station to exit the current service, and the optimal matching robot will exit the service process of guiding the target user.
[0136] In response to receiving a signal to exit the current service, the system determines the target parking area closest to the optimal matching robot within the preset parking areas and sends the parking area location to the optimal matching robot so that the optimal matching robot can proceed to the parking area location. In other words, when the central control console receives a signal to exit the current service, it determines the location of the target parking area closest to the optimal matching robot. In this embodiment, the target parking area is a designated area where robots that are not performing navigation services are parked. After determining the location of the target parking area closest to the optimal matching robot, the optimal matching robot proceeds to the parking area location, waits for the next target user, and provides the navigation service process.
[0137] By determining whether the target user has arrived at the first meeting point before the third control command of the optimally matched robot, the system effectively avoids prolonged ineffective waiting by the robot, allowing it to be reassigned to other tasks and improving robot utilization. When the robot exits its current service, the system finds the nearest target parking area for the robot based on preset parking zones and sends its location information. This not only ensures orderly robot parking but also reduces the time the robot spends searching for parking locations, improving the overall system's operational efficiency.
[0138] In one embodiment of this invention, semantic recognition is performed on the voice information to determine the first target location that the target user needs to reach, including:
[0139] S510: Convert speech information into speech signals;
[0140] S520. Perform endpoint detection on the speech signal to determine the start and end endpoints of the speech signal;
[0141] S530. Obtain the audio segment through the start endpoint and the end endpoint;
[0142] S540. Apply a preset filtering algorithm to the audio segment to obtain the filtered audio segment;
[0143] S550. Input the filtered audio segment into the pre-trained acoustic model to obtain the feature vector of the audio segment;
[0144] S560. Input the filtered audio segment into the pre-trained language model to obtain the word sequence of the audio segment;
[0145] S570. Semantic understanding of audio segments is performed using feature vectors and word sequence to obtain key information about the audio segments;
[0146] S580: Determine the first target location that the target user needs to reach by using key information.
[0147] Semantic recognition of voice information determines the first target location that the target user needs to reach. Semantic recognition is based on artificial intelligence technology, which enables computers to understand and analyze the meaning and semantics of natural language text or speech, and convert the meaning in the text into a computer-readable form. Specifically, firstly, the voice information is converted into a voice signal. This can be done by collecting the target user's sound waves through a microphone on the target terminal, and then converting the collected sound waves into electrical signals to obtain the voice signal.
[0148] After obtaining the speech signal, endpoint detection is performed to determine the start and end points of the speech signal. This can be achieved using a dual-threshold endpoint detection method based on short-time energy and short-time average zero-crossing rate. Short-time energy refers to the speech signal being considered steady-state and invariant within a short time range. Therefore, the speech signal can be divided into different frames to analyze its characteristic parameters; energy is high when speech is present and low when no speech is present. The short-time average zero-crossing rate represents the number of times the waveform of a frame of speech signal crosses the zero level on the horizontal axis. Voiced segments have a low average zero-crossing rate, concentrated in the low-frequency band, while unvoiced segments have a high average zero-crossing rate, concentrated in the high-frequency band. The short-time average zero-crossing rate can be used to find the speech signal from background noise and determine the start and end points of speech segments. Endpoint detection is a crucial step in speech signal processing, mainly used to identify speech and non-speech regions in the speech signal and accurately locate the start and end points of speech. In this embodiment, endpoint detection of speech signals is performed using a threshold-based endpoint detection method. By extracting time-domain features such as short-time energy and short-term zero-crossing rate, or frequency-domain features such as Mel-frequency cepstral coefficients (MFCC) and spectral entropy, and setting a reasonable threshold, the method aims to distinguish between speech and non-speech signals.
[0149] After determining the start and end points of the speech signal, an audio segment can be obtained through the start and end points. In other words, by determining the start and end points of the speech signal, an audio segment containing valid speech content can be extracted from the original audio.
[0150] An audio clip is filtered using a preset filtering algorithm. This algorithm can be based on a Fast Fourier Transform (FFT) method, which converts the time-domain signal to the frequency domain. Filtering is then applied in the frequency domain, such as low-pass, high-pass, or band-pass filtering. Finally, an inverse FFT is performed on the filtered frequency-domain signal to recover the time domain. Filtering audio clips improves audio quality and removes noise. Noise in public places typically comes from air conditioner units and crowds; the preset filtering algorithm can reduce the impact of this noise.
[0151] The filtered audio clip is input into a pre-trained acoustic model to obtain the feature vector of the audio clip. In this embodiment, the feature vector of the audio clip includes features such as pitch, volume, timbre, and rhythm. The feature vector, extracted from the audio signal, reflects important attributes of the audio content. The pre-trained acoustic model can be a convolutional neural network (CNN), a representative algorithm in deep learning, belonging to the category of feedforward neural networks. It has a deep structure and convolutional computation characteristics, extracting local features from the input data through convolution operations to form complex feature representations, which are ultimately used for tasks such as classification and regression. Specifically, the filtered audio clip is input into the pre-trained acoustic model, which outputs the feature vector of the audio clip. The feature vector typically contains information about the audio content, such as pitch, rhythm, and timbre.
[0152] After obtaining the feature vector of the audio segment, the filtered audio segment is input into a pre-trained language model to obtain the word sequence of the audio segment. The word sequence of the audio segment refers to the transformation of an audio signal into a series of words arranged in chronological order after processing by speech recognition technology. The pre-trained language model can be a neural network model, which uses the neural network model to capture the relationship between words and improve the performance of the language model. In other words, the acoustic model is responsible for converting the audio signal into a series of acoustic features such as phonemes, while the language model is used to evaluate the rationality of the sentences or word groups composed of these modeling units.
[0153] By using feature vectors and word sequence sequences, semantic understanding is performed on audio segments to obtain key information. In this embodiment, the key information includes location and time information. An acoustic model converts the speech signal into a pronunciation sequence. Through learning the feature vectors, the acoustic model can identify phonemes and syllables in the speech signal. Based on the acoustic model's results, the word sequence with the highest probability of occurrence is obtained. The language model, by statistically analyzing the relationships between words in the text, can infer the word sequence that best conforms to grammatical rules and semantic logic. Using the word sequence with the highest probability of occurrence, key information about the target user's audio segment is obtained, such as the location the target user needs to reach within a certain time period.
[0154] By using key information, the first target location that the target user needs to reach is determined. That is, by using a pre-trained acoustic model and a pre-trained language model, key information in the target user's audio segment is obtained, and the key information is analyzed. For example, the word sequence that appears most frequently in the key information is determined, and the first target location that the target user needs to reach is determined. In this embodiment, the first target location is the location that the target user needs to reach.
[0155] By accurately identifying the start and end points of the speech signal, silence and noise can be effectively removed, retaining audio segments containing valid speech information. This helps reduce the complexity of subsequent processing and improves recognition accuracy. Through efficient speech processing and semantic recognition algorithms, the system can quickly parse user-input speech information and determine the target location. This enhances the user experience, enabling users to quickly obtain the information or services they need.
[0156] In one embodiment of this invention, endpoint detection is performed on the speech signal to determine the start and end endpoints of the speech signal, including:
[0157] S610. Perform frame segmentation on the speech signal to obtain speech frames of multiple speech signals.
[0158] S620. Perform a Fourier transform on each speech frame to obtain the spectrum of each speech frame.
[0159] S630. Calculate the spectral energy value for each spectrogram;
[0160] S640. When the spectral energy value is less than a preset low threshold and the spectral energy value of the speech frame adjacent to the speech frame whose spectral energy value is less than the preset low threshold is greater than a preset high threshold, the first speech frame whose spectral energy value is less than the preset low threshold is taken as the start endpoint of the speech signal.
[0161] S650. When the spectral energy value is less than a preset high threshold and the spectral energy value of the speech frame adjacent to the speech frame whose spectral energy value is less than the preset high threshold is greater than the preset high threshold, the last speech frame greater than the preset high threshold is taken as the end point of the speech signal.
[0162] Endpoint detection is performed on the speech signal to determine the start and end points. Specifically, the speech signal is first segmented into frames to obtain multiple speech frames. Before segmentation, the speech signal is usually windowed to reduce spectral leakage and truncation effects. Based on the determined frame length, the windowed speech signal is segmented into multiple short segments of the same length as the frame, with some overlap between adjacent frames. Frame shift processing is then applied to this overlap. Frame shift refers to the length of the overlap between adjacent frames during the segmentation process and is an important parameter in speech signal processing to ensure the continuity and smoothness of the speech signal. By segmenting the speech signal, the continuity of the balanced speech signal is guaranteed.
[0163] Secondly, a Fourier transform is performed on each speech frame to obtain its spectrogram. Performing the Fourier transform on each speech frame is a crucial step in speech signal processing; it converts the speech signal from the time domain to the frequency domain, thus obtaining the spectrogram for each speech frame. The spectrogram displays the energy distribution of the speech signal at different frequencies. Applying the Fourier transform to each speech frame, transforming it from the time domain to the frequency domain, yields a complex array where each element corresponds to a frequency component. The amplitude represents the energy level of that frequency component, and the phase represents its phase information. The energy or the square of the amplitude at each frequency component is calculated, and these energy values are plotted to obtain the spectrogram for each speech frame.
[0164] After obtaining the spectrogram of each speech frame, the spectral energy value of each spectrogram is calculated. The spectral energy value is a concept used to describe the amount of energy distributed across different frequency components of a signal. By calculating the spectral energy value of each spectrogram, the start and end points of the speech signal can be determined. The spectral energy value of each spectrogram can be determined by calculating its spectral entropy value. Spectral entropy is a concept in information theory that describes the uncertainty or complexity of the energy distribution of a signal in the frequency domain. A higher spectral entropy value indicates a more uniform energy distribution in the frequency domain, meaning smaller energy differences across frequency components. Conversely, a lower spectral entropy value indicates a more uneven energy distribution in the frequency domain, with significant concentrations of energy in certain frequency components.
[0165] The spectral entropy value of each spectrogram can be calculated using the spectral entropy value calculation formula, as shown below:
[0166] S = -∑p_i*log_2(p_i)
[0167] Where S represents the total entropy value, ∑ represents summation, and p_i represents the energy of each frequency component divided by the total energy, thus obtaining the energy probability of each frequency component;
[0168] The spectral energy value of each spectrogram is obtained by using the formula for calculating spectral entropy. After obtaining the spectral energy value of each spectrogram, it is determined whether the spectral energy value of the spectrogram of each speech frame is less than a preset threshold. The preset threshold can be determined according to the actual situation. That is, the spectral energy value of the spectrogram of each speech frame is compared with the preset threshold.
[0169] When the spectral energy value is less than a preset low threshold and the spectral energy value of the adjacent speech frame is greater than a preset high threshold, the first speech frame with a spectral energy value less than the preset low threshold is taken as the start endpoint of the speech signal. In this embodiment, the preset low threshold and preset high threshold can be determined according to the actual situation. That is, when the spectral energy value is less than the preset low threshold, it means that the energy of the current speech frame in the frequency domain is very low, which may represent a gap or silence in the speech signal. When the spectral energy value of the adjacent speech frame is greater than the preset high threshold, and one or more speech frames following the aforementioned low-energy frames suddenly increase in energy, exceeding the preset high threshold, it usually indicates the start of the speech signal. The speech start endpoint is detected by traversing the speech frame sequence and checking the spectral energy value of each frame. When a frame is found to have a spectral energy value less than the preset low threshold, and the spectral energy value of the next or subsequent frames is greater than the preset high threshold, this frame with a spectral energy value less than the low threshold is marked as the start endpoint of the speech signal.
[0170] When the spectral energy value is less than a preset high threshold and the spectral energy value of the adjacent speech frame is greater than the preset high threshold, the last speech frame with a spectral energy value greater than the preset high threshold is taken as the end point of the speech signal. In other words, when the spectral energy value is less than the preset high threshold, it indicates that the current speech frame has relatively low energy in the frequency domain, possibly meaning that the speech signal is weakening or about to end. When the spectral energy value of an adjacent speech frame is greater than the preset high threshold, there are one or more frames with high energy values before the frames below the high threshold, exceeding the preset high threshold. When a series of low-energy frames appear after these high-energy frames, the speech signal is confirmed to have truly ended. The speech end point is detected by traversing the speech frame sequence, checking the spectral energy value of each frame from beginning to end. When a frame with a spectral energy value less than the preset high threshold is encountered, the energy values of this frame and its subsequent frames are recorded. This process continues until a frame with a spectral energy value greater than the preset high threshold is found again, indicating that there may still be speech signal. If no frame greater than the high threshold is found in subsequent frames, then the last frame less than the high threshold but immediately followed by one or more frames greater than the high threshold, i.e. the last one in this series of low-energy frames, is marked as the end point of the speech signal.
[0171] By performing endpoint detection on the speech signal, the accuracy of starting and ending endpoint detection can be significantly improved, reducing false detections and missed detections. This helps to reduce redundant data in subsequent speech processing and ensures the accuracy of speech signal recognition.
[0172] In one embodiment of this invention, determining the user level of the target user based on the first target location and travel information includes:
[0173] S710. Based on the itinerary information, determine the first time point for the target user;
[0174] S720. Determine the path length from the first current location to the first target location based on the first target location and the first current location;
[0175] S730, Get the current time node;
[0176] S740. Based on the path length and the preset speed of the optimal matching robot, calculate the time period for the optimal matching robot to reach the first target location, and based on the current time node and the time period, calculate the second time node for the optimal matching robot to reach the first target location.
[0177] S750. Map the first path to the preset geographic data model, determine the influencing factors of the first path, and determine the third time node for the optimal matching robot to reach the first target location based on the influencing factors.
[0178] S760. Using the grey relational analysis algorithm, construct an evaluation matrix for the first time node, the second time node, and the third time node;
[0179] S770. Determine the weight values of the first time node, the second time node, and the second time node;
[0180] S780. Using a preset synthesis operator, the weight values and the evaluation matrix are synthesized to obtain the comprehensive evaluation result of the user level.
[0181] S790. Determine the user level of the target user based on the comprehensive evaluation results of the user level.
[0182] Based on the travel information, the first time point of the target user is determined. In this embodiment, the travel information may include multiple aspects such as the user's travel plan, departure time, arrival time, and places passed through. In other words, the first time point of the target user is determined by the travel information of the target user. In this embodiment, the first time point refers to the latest time of the target user's means of transportation. For example, if the target user's means of transportation is a train, the first time point is the latest time point of the train.
[0183] After determining the first time point of the target user, the path length from the first current location to the first target location is determined based on the first target location and the first current location. In other words, the path length from the target user to the first target location is determined by the first target location that the target user needs to reach and the target user's current location. The path length from the target user to the first target location can be determined by indoor positioning technology, such as using wireless communication signals like Wi-Fi and Bluetooth for indoor positioning.
[0184] The system obtains the current time node. Based on the path length and the preset optimal matching robot speed, it calculates the time interval for the optimal matching robot to reach the first target location. In other words, by dividing the previously obtained path length by the preset optimal matching robot speed, the system obtains the preset time interval for the optimal matching robot to reach the first target location. In this embodiment, the preset optimal matching robot speed can be determined based on actual conditions. After obtaining the preset time interval for the optimal matching robot to reach the first target location, the system calculates the second time node for the optimal matching robot to reach the first target location based on the current time node and the time interval. For example, if the current time node is 5:30 PM, and the preset time interval for the optimal matching robot to reach the first target location is calculated to be 30 minutes, then the second time node is 6:00 PM.
[0185] The first path is mapped onto a preset geographic data model, the influencing factors of the first path are determined, and the optimal third time node for the matching robot to reach the first target location is determined based on the influencing factors. In this embodiment, the influencing factors of the first path may be pedestrian traffic and roadblocks. The preset geographic data model can be a Geodatabase model. A Geodatabase is a unified and intelligent spatial database built on top of a standard relational database management system. It is a core component of ArcGIS used to store and manage geospatial data.
[0186] After mapping the first path to a preset geographic data model and determining the influencing factors of the first path, the computer will automatically generate a time prediction formula. This formula will then be displayed on the screen. For example, the time prediction formula generated on the computer might be:
[0187]
[0188] T represents time, L represents path length, and V represents the robot's preset speed.
[0189] After obtaining the time prediction formula, the pedestrian flow and roadblock factors are fitted into the time prediction formula to obtain the adjustment formula:
[0190]
[0191] V represents the robot's preset speed, L represents the path length, V1 represents the formula for the relationship between pedestrian flow and speed, and V2 represents the formula for the relationship between obstacles and speed.
[0192] The formula relating pedestrian flow and speed is:
[0193]
[0194] Q represents the flow of people, and K represents the density of people.
[0195] The formula relating roadblocks and speed is:
[0196] V2=λ a x(l)
[0197] λ a The parameter representing the impact of road obstacles on the robot's speed is x(l), where x(l) represents the number of road obstacles in a path segment of length l.
[0198] Where, λ a The influencing parameters are determined by factors such as the friction, temperature, and humidity of the road surface material. In some cases, humidity affects the coefficient of friction between the robot and the ground or other objects. For example, a wet floor may cause the robot to slip or have difficulty maintaining stability, thus affecting its speed. Furthermore, in high-temperature environments, the robot may require a more powerful cooling system to avoid overheating; in low-temperature environments, the robot may require a heating system to maintain its internal temperature. Therefore, λ varies depending on the road surface. a The values of the influencing parameters differ; for example, the parameters in a road segment are shown in Table 1:
[0199]
[0200] It should be noted that Table 1 is only a partial illustrative example and is not intended to limit the scope of this application.
[0201] In practice, different influencing parameters, such as the material friction, temperature, and humidity of the ground, can be input into the pre-trained model, and the output will be λ corresponding to the material friction, temperature, and humidity of the ground. a Influences the value of the parameter.
[0202] The final adjusted time prediction formula is obtained as follows:
[0203]
[0204] Where L represents the path length, V represents the robot's preset speed, Q represents the number of people, K represents the density of people, and λ a The parameter representing the impact of road obstacles on the robot's speed is x(l), where x(l) represents the number of road obstacles in a path segment of length l.
[0205] After obtaining the adjusted time prediction formula, the third time node of the robot is determined by the adjusted time prediction formula. In this embodiment, the third time node is the time node when the optimal matching robot arrives at the first target location based on the influencing factors. For example, the second time node is 6:00 pm. The adjusted time prediction formula determines that the robot's time extension due to influencing factors is 10 minutes. Therefore, the third time node is 6:10 pm.
[0206] Using the grey relational analysis algorithm, evaluation matrices are constructed for the first, second, and third time nodes. Grey relational analysis is a multivariate statistical analysis method based on grey system theory, primarily used to analyze the degree of correlation between various factors in a system, and can be used to construct evaluation matrices. An evaluation matrix is a tool used to assess and analyze the relationships between multiple variables or factors, typically presented in a structured manner to facilitate comparison and quantification of the importance, influence, or performance of different factors. In other words, by constructing evaluation matrices using the grey relational analysis algorithm, the different levels of importance of the first, second, and third time nodes to the target user's time nodes can be determined, thereby identifying the influencing factors of each time node. By analyzing these influencing factors, the urgency level of the target user at each node can be obtained, thus determining the target user's user level.
[0207] The weights of the first, second, and third time nodes can be determined using the Analytic Hierarchy Process (AHP). AHP is a structured technique for complex decision-making problems and a decision analysis method that combines qualitative and quantitative approaches. Mathematical methods such as eigenvector method, arithmetic mean method, and geometric mean method are used to calculate the eigenvalues and eigenvectors of the judgment matrix, and the eigenvectors are normalized to obtain the weights of each time node and related factors.
[0208] After determining the weight values, a preset synthesis operator is used to synthesize the weight values and the evaluation matrix to obtain the comprehensive user level evaluation result. In this embodiment, the preset synthesis operator can be a geometric mean type. The geometric mean type refers to the nth root of the product of n observations. The geometric mean type is a calculation method for the average of a set of values. The weight values of the first time node, the second time node, and the second time node are multiplied with the evaluation matrix using the geometric mean type to obtain the evaluation result vector. Through the evaluation result vector, the membership degree of the evaluated object on each comment is further analyzed to obtain the comprehensive user level evaluation result. For example, the evaluation result vector is B = (b1, b2, b3, ... b... mThe evaluation result vector B is normalized using a linear function, which transforms the original data linearly to the range of [0,1] or [-1,1] to obtain the normalized data and determine the comprehensive evaluation result of the user level.
[0209] The user level of the target user is determined by the comprehensive evaluation results of the user level. In other words, within the preset level evaluation threshold range, the comprehensive evaluation results of the user level are compared with the preset level evaluation threshold range to determine the level evaluation threshold range in which the target user is located, thereby determining the user level of the target user.
[0210] By identifying the user levels of target users, it is possible to more accurately assess the urgency and needs of users. Based on the user levels, service providers can allocate resources more effectively, which helps to improve service efficiency, ensure that resources are used rationally, and enhance the user experience.
[0211] In one embodiment of this invention, the robot information includes robot device information. Determining the optimal matching robot based on the user level, the first target location, and the robot information includes:
[0212] S810: Obtain spatial location data of the robot;
[0213] S820: Map the target user, primary target location, and robot to the geographic data model and display them;
[0214] S830. Based on the user level, determine the user's time threshold in the preset user level database. The time threshold includes a first time period and a second time period. The first time period is the time period between the first current location and the first target location, and the second time period is the time period between the second current location and the first current location.
[0215] S840. Retrieve the target robot in the geographic data model by using a time threshold;
[0216] S850. Determine environmental information of public places through geographic data models;
[0217] S860: By using environmental information from public places, determine in real time the time required for the target robot to reach the first target location by passing through the location of the target user;
[0218] S870. Determine the electrical energy reserve of the target robot based on the robot equipment information of the target robot;
[0219] S880. When the target robot's energy reserves are greater than the pre-consumed energy and the time required for the target robot to reach the first target location via the first current location is less than a preset time threshold, determine the optimal matching robot.
[0220] Based on the user level, the primary target location, and robot information, the optimal matching robot is determined. Specifically, firstly, the spatial location data of the robot is obtained. This can be achieved through LiDAR scanning or on-site measurement, such as the boundaries and extent of the space, the shape of buildings, and the width of roads.
[0221] Subsequently, the target user, the first target location, and the robot are mapped to a preset geographic data model and displayed. That is, the latitude and longitude coordinates of the target user are converted into the geographic data model, and the current location information of the robot is converted into the geographic data model and displayed in a visualization tool or platform, such as GIS software, map services, or custom map visualization applications.
[0222] Based on the user's level, a time threshold for the user is determined in a preset user level database. In other words, the time threshold corresponding to the target user is determined by the user's level in the preset user level database. This time threshold includes a first time period and a second time period. The first time period is the time required for the target user to travel to the first target location, and the second time period is the time from the location of the optimal matching robot to the target user. By determining these two time periods and the user's level, the user's time threshold is determined. For example, if the target user's latest travel time is 6:00, and there is one hour remaining before the latest travel time, the time required for the target user to travel to the first target location is 30 minutes, and the time from the location of the optimal matching robot to the target user is 20 minutes. Simultaneously, an extended time period is determined for the user in the preset user level database to allow time for the user to encounter other inconveniences during the journey, thus ultimately determining the user's time threshold.
[0223] By using time thresholds, the target robot is retrieved from the geographic data model. In other words, based on the time it takes for the optimal matching robot to locate the target user and the user's time threshold, the geographic data model determines the robot whose time to locate the target user is less than the user's time threshold, thus obtaining the target robot.
[0224] By using geographic data models to determine the environmental information of public places, GIS software can be used for spatial analysis, such as buffer analysis and overlay analysis, to determine the spatial distribution characteristics of indoor public places. Spatial analysis can also be used to assess environmental indicators such as accessibility and congestion of indoor public places.
[0225] After determining the environmental information of the public place, the time required for the target robot to reach the first target location via the target user's location and the estimated power consumption are determined in real time. Specifically, firstly, the precise coordinates of the target user's location and the first target location need to be obtained. Using map data or an indoor navigation system, the optimal path for the robot from the target user's location to the first target location is planned. Based on the target robot's preset speed, the time for the target robot to reach the first target location is calculated. Subsequently, the target robot's battery capacity and energy consumption rate, as well as environmental information, such as the impact of environmental factors like temperature and humidity on battery performance, are obtained. Based on the path length, the target robot's moving speed, energy consumption rate, and environmental information, the power consumption of the target robot along the entire path is calculated. Empirical formulas or simulation models can be used to estimate energy consumption.
[0226] After obtaining the time and pre-consumed power required for the target robot to reach the first target location via the target user's location, the power reserve of the target robot is determined through the robot's equipment information. In other words, the power reserve is determined based on the target robot's equipment information. When the target robot's power reserve is greater than the pre-consumed power and the time required for the target robot to reach the first target location via the current location is less than a preset time threshold, the optimal matching robot is determined. At this time, the target robot's power reserve is sufficient to offset the pre-consumed power, and the target robot also has enough time to guide the target user to the first target location.
[0227] By comprehensively considering user level, target location, and robot information, the most suitable robot can be accurately matched to the user, which can improve the efficiency of robot use, thereby ensuring service quality and improving user satisfaction and experience.
[0228] In one embodiment of this invention, after determining the first target location that the target user needs to reach and performing facial recognition on the facial image to obtain the target user's travel information, the method further includes:
[0229] S910. Retrieve the second target location closest to the target user within the geographic data model;
[0230] S920. Based on the target user's travel information, obtain the target user's destination and determine whether the first target location and the second target location are the destinations;
[0231] S930. If both the first target location and the second target location are destinations, generate second target location confirmation information and send it to the target terminal based on the second target location.
[0232] S940. After receiving the confirmation information sent by the target terminal, the second target location is designated as the first target location.
[0233] S950. After receiving the negative information sent by the target terminal, perform the steps of real-time acquisition of robot information for each robot and real-time acquisition of the first current location of the target user.
[0234] Figure 2 A schematic diagram illustrating a target location determination process provided in an embodiment of this application;
[0235] The second target location closest to the target user is retrieved within the geographic data model. In other words, the current location of the target user is determined, and the second target location closest to the target user is retrieved within the geographic data model. In this embodiment, the second target location and the first target location are both path locations that can lead to the final destination. For example, if the means of transportation taken by the target user has two security checkpoints, the second target location retrieved within the geographic data model is the one closest to the target user's security checkpoint.
[0236] After retrieving the second target location, the system determines the target user's destination based on their travel information and then checks whether the first and second target locations are indeed the destination. For example, the travel information can be used to determine the mode of transportation the target user used and the security checkpoints where security checks can be conducted. Determining whether the first and second target locations are indeed the destination means determining whether they are security checkpoints where security checks can be conducted.
[0237] If both the first and second target locations are destinations, a second target location confirmation message is generated and sent to the target terminal based on the second target location. If both the first and second target locations are security checkpoints where security checks can be conducted, a second target location confirmation message is generated and sent to the target terminal based on the second target location closest to the target user. The target user needs to determine whether to replace the second target location with the first target location.
[0238] Once the target user decides to replace the first target location with the second target location, after receiving confirmation from the target terminal, the second target location will be used as the first target location. In other words, the target user agrees to undergo security checks at the nearest security checkpoint.
[0239] When the target user does not agree to replace the second target location with the first target location, after receiving the negative information sent by the target terminal, the steps of obtaining the robot information of each robot in real time and obtaining the first current location of the target user in real time will be executed. In other words, when the target user does not agree to replace the second target location with the first target location, for example, when the target user needs to go to the first target location to handle other matters, the steps of obtaining the robot information of each robot in real time and obtaining the first current location of the target user in real time will be executed.
[0240] In another implementation of this embodiment, for an airport check-in scenario, after obtaining the target user's destination through their travel information, if it is found that the first target location (first security checkpoint) the target user wants to reach cannot be used for security checks, while the second target location (second security checkpoint) can be used, a confirmation message is sent to the target user. The message content could be: "Second security checkpoint can be used for security checks, while first security checkpoint cannot be used for security checks. Do you want to continue to the first security checkpoint?". Upon receiving the target user's instruction to go to the second target location, the first target location the target user wants to reach is corrected, replacing it with the second target location. After receiving the instruction from the target user to go to the first target location, the process continues to execute the steps of real-time acquisition of robot information for each robot, and simultaneously real-time acquisition of the target user's first current location. This application will not elaborate further on this aspect.
[0241] By introducing a second destination location retrieval and confirmation mechanism, we can better understand user intent and provide services that better meet user needs.
[0242] This application embodiment also provides an interactive robot, each interactive robot being connected to a central control console, and each interactive robot being equipped with a first camera. The interactive robot further includes:
[0243] The memory is configured to store instructions; and
[0244] The processor is configured to retrieve instructions from memory and execute them.
[0245] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0246] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0247] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0248] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0249] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0250] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0251] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0252] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0253] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A control method for an intelligent interactive robot in public places, characterized in that, The method, applied to a central control console connected to at least one robot, wherein each robot is equipped with a first camera, includes: In response to receiving path navigation instructions from multiple user terminals, the user terminal corresponding to the earliest path navigation instruction is taken as the target terminal. In response to receiving voice information sent by the target terminal, a shooting command is sent to the target terminal. The shooting command is used to activate the second camera of the target terminal so that the target terminal can capture a facial image of the target user corresponding to the target terminal. In response to receiving the face image sent by the target terminal, semantic recognition is performed on the voice information to determine the first target location that the target user needs to reach, and face recognition is performed on the face image to obtain the travel information of the target user; The robot information of each robot is acquired in real time, and the first current location of the target user is acquired in real time. Based on the first target location and the travel information, the user level of the target user is determined; Based on the user level, the first target location, each robot information and the first current location, determine the optimal matching robot and obtain the second current location of the optimal matching robot; Based on the first target location, the first current location, and the second current location, generate and display the optimal matching robot's path control scheme; A first path is generated based on the second current location, the first current location, and the first target location; The step of determining the user level of the target user based on the first target location and the travel information includes: Based on the travel information, the first time point of the target user is determined; Based on the first target location and the first current location, determine the path length from the first current location to the first target location; Get the current time point; Based on the path length and the preset speed of the optimal matching robot, calculate the time period for the optimal matching robot to reach the first target location, and based on the current time node and the time period, calculate the second time node for the optimal matching robot to reach the first target location. The first path is mapped into a preset geographic data model, the influencing factors of the first path are determined, and the third time node for the optimal matching robot to reach the first target location is determined based on the influencing factors. Using the grey relational analysis algorithm, an evaluation matrix is constructed from the first time node, the second time node, and the third time node; Determine the weight values of the first time node, the second time node, and the second time node; The weight values and the evaluation matrix are combined using a preset synthesis operator to obtain a comprehensive user level evaluation result. The user level of the target user is determined based on the comprehensive evaluation results of the user level.
2. The method according to claim 1, characterized in that, The step of generating the path control scheme for the optimal matching robot based on the first target location, the first current location, and the second current location includes: The first path includes the second current location, the first current location, and the first target location. The second current location is the first waypoint of the first path, the first current location is the second waypoint of the first path, and the first target location is the third waypoint of the first path. Based on the first path, a first control command is generated and sent to the optimal matching robot so that the optimal matching robot executes the first control command.
3. The method according to claim 2, characterized in that, The method further includes: A first rendezvous point is generated based on the second current location and the first target location; A second path is generated using the first current location and the first meeting point. The second path includes the first current location and the first meeting point. The first current location is the first waypoint of the second path, and the first meeting point is the second waypoint of the second path. Send the second path to the target terminal; A third path is generated using the first rendezvous point and the second current location. The third path includes the first rendezvous point and the second current location. The second current location is the first waypoint of the third path, and the first rendezvous point is the second waypoint of the third path. Based on the third path, a second control command is generated and sent to the optimal matching robot so that the optimal matching robot executes the second control command; A fourth path is generated using the first meeting point and the first target location. The fourth path includes the first meeting point and the first target location. The first meeting point is the first waypoint of the fourth path, and the first target location is the second waypoint of the fourth path. Based on the fourth path, a third control command is generated and sent to the optimal matching robot so that the optimal matching robot executes the third control command.
4. The method according to claim 3, characterized in that, Before generating the third control command according to the fourth path, the process includes: The arrival time of the target user is determined using the second path; The preset waiting time threshold for the optimal matching robot is set based on the arrival time; Within a preset waiting time threshold, the first camera of the optimal matching robot is controlled to capture images; The similarity between the image and the face image is calculated. If the similarity is greater than a preset similarity threshold, the optimal matching robot executes a third control command. If the waiting time of the optimal matching robot is greater than a preset waiting time threshold and the similarity is less than a preset similarity threshold, then the optimal matching robot sends a signal to the central control console to exit the current service. In response to receiving the signal to exit the current service, a target parking area closest to the optimal matching robot is determined in the preset parking area, and the location of the parking area is sent to the optimal matching robot so that the optimal matching robot can go to the parking area location.
5. The method according to claim 1, characterized in that, The step of performing semantic recognition on the voice information to determine the first target location that the target user needs to reach includes: The voice information is converted into a voice signal; Endpoint detection is performed on the speech signal to determine the start and end endpoints of the speech signal; An audio segment is obtained through the start endpoint and the end endpoint; The audio segment is processed using a preset filtering algorithm to obtain a filtered audio segment; The filtered audio segment is input into a pre-trained acoustic model to obtain the feature vector of the audio segment; The filtered audio segment is input into a pre-trained language model to obtain a word sequence of the audio segment; By using the feature vector and the word sequence, semantic understanding is performed on the audio segment to obtain the key information of the audio segment; Based on the key information, the first target location that the target user needs to reach is determined.
6. The method according to claim 5, characterized in that, The endpoint detection of the speech signal to determine the start and end endpoints of the speech signal includes: The speech signal is segmented into frames to obtain multiple speech frames of the speech signal; Perform a Fourier transform on each of the speech frames to obtain the spectrogram of each speech frame; Calculate the spectral energy value for each of the aforementioned spectrograms; When the spectral energy value is less than a preset low threshold and the spectral energy value of the speech frame adjacent to the speech frame whose spectral energy value is less than the preset low threshold is greater than a preset high threshold, the first speech frame whose spectral energy value is less than the preset low threshold is taken as the start endpoint of the speech signal. When the spectral energy value is less than a preset high threshold and the spectral energy value of the speech frame adjacent to the speech frame whose spectral energy value is less than the preset high threshold is greater than the preset high threshold, the last speech frame whose spectral energy value is greater than the preset high threshold is taken as the end point of the speech signal.
7. The method according to claim 1, characterized in that, The robot information includes robot device information. The step of determining the optimal matching robot based on the user level, the first target location, and the robot information includes: Obtain spatial location data of the robot; The target user, the first target location, and the robot are mapped into the geographic data model and displayed. Based on the user level, a time threshold for the user is determined in a preset user level database. The time threshold includes a first time period and a second time period. The first time period is the time period between the first current location and the first target location, and the second time period is the time period between the second current location and the first current location. The target robot is retrieved from the geographic data model using the time threshold. The environmental information of the public place is determined using the geographic data model. Based on the environmental information of the public place, the time required for the target robot to reach the first target location by passing through the location of the target user and the pre-consumed power can be determined in real time. The electrical energy reserve of the target robot is determined by the robot equipment information of the target robot. When the target robot's energy reserves are greater than the pre-consumed energy and the time required for the target robot to reach the first target location via the first current location is less than a preset time threshold, the optimal matching robot is determined.
8. The method according to claim 7, characterized in that, After determining the first target location that the target user needs to reach, and performing facial recognition on the facial image to obtain the target user's travel information, the method further includes: Retrieve the second target location closest to the target user within the geographic data model; Based on the target user's travel information, the target user's destination is obtained, and it is determined whether the first target location and the second target location are the destination; If both the first target location and the second target location are the destination, then a second target location confirmation message is generated and sent to the target terminal based on the second target location. After receiving the confirmation information sent by the target terminal, the second target location is taken as the first target location; Upon receiving the negative information sent by the target terminal, the steps of acquiring robot information of each robot in real time and acquiring the first current location of the target user in real time are executed.
9. An interactive robot, applied to the robot according to any one of claims 1-8, characterized in that, Each of the interactive robots is connected to a central control panel, and each of the interactive robots is equipped with a first camera. The interactive robot also includes: The memory is configured to store instructions; and The processor is configured to retrieve and execute the instructions from the memory.
Citation Information
Patent Citations
Unmanned intelligent perception service system
CN115759493A
Intelligent carrying method and related device
CN115848408A