Visual inspection-based robotic automated survey method

By using a robot camera to screen for target individuals who are easy to survey during two rounds of panoramic viewing, the problem of inaccurate selection of survey subjects by robots is solved, thus improving the quality and success rate of surveys.

CN121169899BActive Publication Date: 2026-04-21上海万怡医学科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
上海万怡医学科技股份有限公司
Filing Date
2025-09-29
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

When robots conduct surveys in venues such as conference rooms, they may select inaccurate survey subjects, leading to a decline in the quality of the surveys.

Method used

The robot's camera captures images from multiple angles during two rounds of surround view. Using image detection and neural network models, it identifies target researchers who are likely to participate in the survey and then moves the robot to them to extend a survey invitation.

Benefits of technology

This improved the accuracy of researcher selection and the success rate of research, thereby enhancing the quality of the research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121169899B_ABST
    Figure CN121169899B_ABST
Patent Text Reader

Abstract

This invention provides a visual detection-based automated robot survey method, relating to the field of robotics. The method includes: capturing first and second images from multiple angles and screening potential survey participants; obtaining a first region to be analyzed within the first image and a second region to be analyzed within the second image for each potential survey participant, and determining a survey convenience index to screen target survey participants, thereby sending survey invitations to the target participants. According to this invention, potential survey participants can be screened using first and second images captured by the robot's camera during two rounds of surround view. The method automatically analyzes whether the region where the participant is located is convenient for surveying, thereby reducing the randomness of participant selection and automatically selecting participants who are convenient for surveying, thus improving survey quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics, and more particularly to a visual inspection-based automated surveying method for robots. Background Technology

[0002] During research activities conducted in venues such as conference rooms, robots can be used to process the research. For example, a mobile robot can move near a person and display a survey questionnaire or distribute a paper survey questionnaire to that person. However, the survey respondents are usually randomly selected or selected from the nearest available person. But randomly selected or nearby respondents may not be convenient to receive the survey. For example, if the selected respondents are busy with other things, they may refuse the survey or answer the survey questionnaire perfunctorily, resulting in a decline in the quality of the survey. Summary of the Invention

[0003] This invention provides a visual inspection-based automatic survey method for robots, which can solve the technical problem of inaccurate selection of survey subjects by robots, leading to a decline in survey quality in related technologies.

[0004] According to a first aspect of the present invention, a vision-based automatic surveying method for robots is provided, comprising:

[0005] The robot's camera captures first images from multiple angles during the initial surround view process;

[0006] The robot's camera captures second images from multiple angles during the second round of surround view.

[0007] Based on the first image and the second image, select potential research subjects from among multiple people around the robot;

[0008] Obtain the first region to be analyzed in the first image where the researchers to be identified are located.

[0009] Obtain the second region to be analyzed in the second image where the researchers to be identified are located;

[0010] Based on the first region to be analyzed and the second region to be analyzed, determine the survey convenience index for the candidates to be surveyed;

[0011] Based on the aforementioned survey convenience index, target survey personnel are selected from the pool of undetermined survey personnel;

[0012] Control the robot's movement components to move the robot to a preset distance in front of the target researcher's face and send out a research invitation message.

[0013] According to the present invention, based on the first image and the second image, selecting a candidate for investigation from among multiple people surrounding the robot includes:

[0014] Detection is performed on the first image from multiple angles to obtain the first region where the person is located in each first image, wherein the first region is a rectangular area that selects the range of the person's location;

[0015] By using an image detection model, feature extraction processing is performed on the first region to obtain the first appearance feature information of each person;

[0016] Detection is performed on the second images from multiple angles to obtain the second region where the person is located in each second image;

[0017] By using an image detection model, feature extraction processing is performed on the second region to obtain the second appearance feature information of each person;

[0018] Based on the first and second facial feature information, a target person is identified that exists in both the first and second images.

[0019] Obtain the first target region of the target person in the first image and the second target region in the second image;

[0020] Based on the first and second target areas, select potential research personnel from among the target personnel.

[0021] According to the present invention, screening potential research subjects from target personnel based on a first target area and a second target area includes:

[0022] Determine the first centroid coordinate data of the first target region and the second centroid coordinate data of the second target region;

[0023] Based on the calibration parameters of the robot's camera, determine the first estimated coordinate data of the first centroid coordinate data in a three-dimensional coordinate system with the location of the camera as the origin, and the second estimated coordinate data of the second centroid coordinate data in the three-dimensional coordinate system.

[0024] According to the formula

[0025]

[0026] Obtain the filtering conditions C1 and C2, where, For the second estimated coordinate data of the i-th target person, The first estimated coordinate data for the i-th target person. For the preset angle threshold, The preset distance threshold;

[0027] Individuals who meet at least one of the screening criteria C1 and C2 are identified as potential survey participants.

[0028] According to the present invention, based on the first region to be analyzed and the second region to be analyzed, the survey convenience index of the candidates to be surveyed is determined, including:

[0029] The first region to be analyzed is divided into a first head region, a first upper body region, and a first lower body region in the vertical direction.

[0030] The second region to be analyzed is divided vertically into a second head region, a second upper body region, and a second lower body region.

[0031] The first head region is processed using the first convolutional neural network model to obtain the first head feature vector, and the second head region is processed to obtain the second head feature vector.

[0032] The first head feature vector and the second head feature vector are concatenated to obtain head feature information;

[0033] The head feature information is input into the first multilayer perceptron model to obtain the head feature description vector.

[0034] The first upper body region is processed using a second convolutional neural network model to obtain the first upper body feature vector, and the second upper body region is processed to obtain the second upper body feature vector.

[0035] The first upper body feature vector and the second upper body feature vector are concatenated to obtain upper body feature information;

[0036] The upper body feature information is input into the second multilayer perceptron model to obtain the upper body feature description vector;

[0037] The first lower body region is processed using a third convolutional neural network model to obtain the first lower body feature vector, and the second lower body region is processed to obtain the second lower body feature vector.

[0038] The first lower body feature vector and the second lower body feature vector are concatenated to obtain lower body feature information;

[0039] The lower body feature information is input into the third multilayer perceptron model to obtain the lower body feature description vector;

[0040] Based on the head feature description vector, upper body feature description vector, and lower body feature description vector, obtain the behavioral category information of the researchers to be determined;

[0041] Based on the behavioral category information, the first region to be analyzed, and the second region to be analyzed, the survey convenience index is determined.

[0042] According to the present invention, behavioral category information of the survey participants is obtained based on head feature description vectors, upper body feature description vectors, and lower body feature description vectors, including:

[0043] The head feature description vector, upper body feature description vector, and lower body feature description vector are used as the input vectors of the nodes in the graph structure, respectively.

[0044] By using the attention mechanism of the graph neural network model, the input vectors of each node are processed to obtain the connection weights between nodes.

[0045] The connection weights and input vectors are processed by a graph neural network model to obtain head feature output vectors, upper body feature output vectors, and lower body feature output vectors.

[0046] The head feature output vector, upper body feature output vector, and lower body feature output vector are concatenated to obtain behavioral description information;

[0047] The behavioral description information is processed by the fourth multilayer perceptual network model to obtain the behavioral category information of the researchers to be surveyed.

[0048] According to the present invention, the method further includes:

[0049] The first training region of the trainee in the first training image is divided into the first training head region, the first training upper body region, and the first training lower body region.

[0050] The second training region in the second training image is divided into the second training head region, the second training upper body region, and the second training lower body region.

[0051] Based on the first training head region and the second training head region, obtain the head training feature vector; based on the first training upper body region and the second training upper body region, obtain the upper body training feature vector; based on the first training lower body region and the second training lower body region, obtain the lower body training feature vector.

[0052] Based on the head training feature vector, upper body training feature vector, and lower body training feature vector, obtain the head training output vector, upper body training output vector, and lower body training output vector.

[0053] The head training output vector is input into the first fully connected layer and the first activation layer to obtain the predicted head action category information;

[0054] The upper body training output vector is input into the second fully connected layer and the second activation layer to obtain the predicted upper body movement category information;

[0055] The lower body training output vector is input into the third fully connected layer and the third activation layer to obtain the predicted lower body movement category information;

[0056] The head training output vector, upper body training output vector, and lower body training output vector are concatenated to obtain behavioral training information.

[0057] The behavioral training information is input into the fourth multilayer perceptron model to obtain the predicted behavioral category information.

[0058] The loss function is determined based on the predicted behavior category information, predicted head movement category information, predicted upper body movement category information, predicted lower body movement category information, and the annotation information of the trainees.

[0059] The loss function is used to train a first convolutional neural network model, a first multilayer perceptron model, a second convolutional neural network model, a second multilayer perceptron model, a third convolutional neural network model, a third multilayer perceptron model, a graph neural network model, a fourth multilayer perceptron model, a first fully connected layer and a first activation layer, a second fully connected layer and a second activation layer, and a third fully connected layer and a third activation layer.

[0060] According to the present invention, a loss function is determined based on predicted behavior category information, predicted head movement category information, predicted upper body movement category information, predicted lower body movement category information, and the annotation information of the trainee, including:

[0061] According to the formula

[0062]

[0063] Determine the loss function LOSS, where, Let be the probability that the trainee's head movement, determined based on the annotation information, is the k-th type of movement. Let be the probability that a trainee's head movement is the k-th type, determined based on the predicted head movement category information. Let be the probability that the trainee's upper body movement, determined based on the annotation information, is the t-th type of movement. Let be the probability that a trainee's upper body movement is the t-th type, determined based on the predicted upper body movement category information. Let s be the probability that the trainee's lower body movement is the s-th type, determined based on the annotation information. Let s be the probability that a trainee's lower body movement is the s-th type, determined based on the predicted lower body movement category information. Let $\mathbf{j}$ be the probability that the behavior category of the trainee, determined based on the annotation information, belongs to the $j$ category. Let $\mathbf{j}$ be the probability that the trainee's behavior category is the $j$-th category, determined based on the predicted behavior category information. The number of categories of head movements. The number of categories of upper body movements. The number of categories of lower body movements. For the number of behavior categories, , , and For the preset weights, k≤ , t≤ s≤ j≤ And k, ,t, ,s, j All are positive integers.

[0064] According to the present invention, a survey convenience index is determined based on the behavioral category information, the first region to be analyzed, and the second region to be analyzed, including:

[0065] Based on the first centroid coordinate data of the first region to be analyzed, the first estimated coordinate data of the person to be investigated in the three-dimensional coordinate system with the location of the camera as the origin is determined, and the second estimated coordinate data of the second centroid coordinate of the second region to be analyzed in the three-dimensional coordinate system is determined.

[0066] Based on the first and second estimated coordinate data, the predicted motion trajectory and motion speed data of the personnel to be investigated in the three-dimensional coordinate system are determined.

[0067] Based on the behavior category information, the predicted motion trajectory data, the motion speed data, and the robot's motion speed, a survey convenience index is determined.

[0068] According to the present invention, a survey convenience index is determined based on the behavior category information, the predicted motion trajectory data, the motion speed data, and the robot's motion speed, including:

[0069] According to the formula

[0070]

[0071]

[0072]

[0073]

[0074] Determine the survey convenience index for the g-th undetermined survey participant. ,in, Let be the probability that the behavior category determined based on the behavior category information of the g-th undetermined survey participant is the j-th category. For the number of behavior categories, The preset convenience coefficient for the behavior of the j-th category. The time required for the robot to meet with the g-th undetermined researcher. For the g-th undetermined survey participant, the second estimated coordinate data represents the planar coordinates. For the g-th undetermined survey personnel, the first estimated coordinate data represents the planar coordinates. Let g be the angle of motion of the g-th undetermined researcher. For the movement speed data of the g-th undetermined researcher, , The time difference between the moment when the undetermined research personnel were captured during the first round of the surround view and the moment when they were captured during the second round of the surround view. For the robot's movement speed, Let j be the robot's direction angle, j≤ And j and All are positive integers.

[0075] According to a second aspect of the present invention, a vision-based automated surveying system for robots is provided, comprising:

[0076] The first image module is used to capture first images from multiple angles during the first round of surround view using the robot's camera;

[0077] The second image module is used to capture second images from multiple angles during the second round of surround view using the robot's camera;

[0078] A screening module is used to select potential survey participants from multiple people around the robot based on the first image and the second image.

[0079] The first region to be analyzed module is used to obtain the first region to be analyzed in the first image where the person to be surveyed is located.

[0080] The second region to be analyzed module is used to obtain the second region to be analyzed in the second image where the person to be surveyed is located.

[0081] The survey convenience index module is used to determine the survey convenience index of the candidates based on the first region to be analyzed and the second region to be analyzed.

[0082] The target survey personnel module is used to select target survey personnel from the undetermined survey personnel based on the survey convenience index.

[0083] The survey invitation module is used to control the robot's movement components, causing the robot to move to a preset distance in front of the target survey participant's face and send a survey invitation message.

[0084] By adopting the above technical solution, the present invention can achieve the following technical effects:

[0085] According to the present invention, candidates for research can be screened using first and second images captured by the robot's camera during two rounds of surround view. The system automatically analyzes whether a candidate is suitable for research based on their location, thereby reducing the randomness of candidate selection and automatically selecting those who are more likely to be surveyed, thus improving research quality. When screening candidates, the robot can estimate their motion by estimating their coordinates in a unified three-dimensional coordinate system, and then set screening conditions to identify candidates moving towards the robot or moving at a slower speed. This facilitates the robot's convergence with the candidate for research and improves the accuracy of candidate screening, thus increasing the success rate of the research. When determining behavioral category information, the robot can divide the first and second analysis areas and process the identical parts after division to obtain feature description vectors describing the actions and postures of each part. A graph neural network model can then be used to obtain the correlation between the actions and postures of each part, thereby improving the accuracy of the overall behavioral category information of the candidates while accurately identifying the actions and postures of each part. When training neural network models, the head, upper body, and lower body training output vectors can be optimized separately. By combining the correlations between these parts, the accuracy of motion prediction for each part can be improved, thereby enhancing the accuracy of the head, upper body, and lower body training output vectors. This increases training intensity and provides a more accurate data foundation for obtaining more precise predicted behavior category information. Furthermore, the predicted behavior category information can be optimized holistically to train multiple neural network models as a whole, further increasing training intensity and improving the accuracy of behavior category prediction. When determining survey convenience indicators, the predicted motion trajectory and speed data of the candidates can be set, and the time required for the candidates to meet with the robot can be calculated. Then, the survey convenience indicators can be determined by using behavior category information and the required meeting time. This allows for the selection of candidates who are easy to meet with the robot and convenient to accept the survey, improving the survey's efficiency and quality. Attached Figure Description

[0086] Figure 1 An exemplary flowchart of a vision-based automated robot survey method according to an embodiment of the present invention is shown.

[0087] Figure 2 A schematic diagram illustrating the determination of behavior category information according to an embodiment of the present invention is shown as an example.

[0088] Figure 3A block diagram of a vision-based robotic automated survey system according to an embodiment of the present invention is shown as an example. Detailed Implementation

[0089] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0090] Figure 1 An exemplary flowchart of a vision-based automated surveying method for robots according to an embodiment of the present invention is shown, the method comprising:

[0091] Step S1: During the first round of surround view, the robot's camera captures first images from multiple angles.

[0092] Step S2: During the second round of surround view, the robot's camera captures second images from multiple angles.

[0093] Step S3: Based on the first image and the second image, select the candidate for investigation from among the multiple people around the robot;

[0094] Step S4: Obtain the first region to be analyzed in the first image where the candidate for research is located;

[0095] Step S5: Obtain the second region to be analyzed in the second image where the candidate researcher is located;

[0096] Step S6: Determine the survey convenience index for the candidates based on the first region to be analyzed and the second region to be analyzed;

[0097] Step S7: Select target researchers from the pool of undetermined researchers based on the survey convenience index.

[0098] Step S8: Control the robot's movement components to move the robot to a preset distance in front of the target researcher's face and send out a research invitation message.

[0099] According to an embodiment of the present invention, the robot-based automatic survey method based on vision detection can screen potential survey personnel by using the first and second images captured by the robot's camera during two rounds of surround view, and automatically analyze whether the survey personnel are convenient to be surveyed based on their location, thereby reducing the randomness of the selection of survey personnel and automatically selecting personnel who are convenient to be surveyed, which helps to improve the quality of the survey.

[0100] According to one embodiment of the present invention, in step S1, the robot's camera performs a first round of panoramic viewing, that is, the yaw angle of the camera changes from 0° to 360°, allowing the camera to perform one panoramic view, and in this process, captures first images from multiple angles. For example, a first image is captured every 45° change in the yaw angle. Similarly, in step S2, during the second round of panoramic viewing, the robot's camera captures multiple second images, and the first and second images with the same sequence number have the same yaw angle, that is, the same field of view.

[0101] According to one embodiment of the present invention, in step S3, the first image and the second image with the same yaw angle were captured at different times. Therefore, the people present in the first image and the second image may change, and their positions and postures may also change. Personnel to be surveyed can be selected based on the first image and the second image. For example, personnel with whom the robot can rendezvous can be selected. In this example, by analyzing the first image and the second image, it can be found that some people are rapidly moving away from the robot, making it difficult for the robot to catch up with them, rendezvous with them, and conduct surveys on them. Therefore, these personnel can be excluded, and personnel with whom the robot can rendezvous can be selected as personnel to be surveyed.

[0102] According to an embodiment of the present invention, selecting potential survey subjects from multiple people around a robot based on the first image and the second image includes: detecting the first image from multiple angles to obtain a first region where a person is located in each first image, wherein the first region is a rectangular area that encloses the area where the person is located; performing feature extraction processing on the first region using an image detection model to obtain first facial feature information of each person; detecting the second image from multiple angles to obtain a second region where a person is located in each second image; performing feature extraction processing on the second region using an image detection model to obtain second facial feature information of each person; determining target subjects that exist in both the first and second images based on the first and second facial feature information; obtaining a first target region of the target subject in the first image and a second target region in the second image; and selecting potential survey subjects from the target subjects based on the first and second target regions.

[0103] According to one embodiment of the present invention, a neural network model such as the YOLO model can be used to detect the first image, obtaining the first region where each person in the first image is located. A bounding box can be used to select the first region. Furthermore, an image detection model can be used to extract features from the first region; for example, the ResNet model can be used as the image detection model to extract feature information from each first region, obtaining the first facial feature information of each person. Similarly, the second image can be processed in the same way to obtain the second facial feature information of each person in the second image.

[0104] According to one embodiment of the present invention, for the first facial feature information of each person in the first image, it is possible to search for whether there is a second facial feature information among a plurality of second facial feature information that has a similarity (e.g., cosine similarity) higher than or equal to a preset similarity threshold (e.g., 0.8) with the first facial feature information. If there is, the first facial feature information and the second facial feature information with a similarity higher than or equal to the preset similarity threshold are determined as the facial feature information of the person present in both the first image and the second image. This person is the target person, and thus the first target region of the target person in the first image and the second target region in the second image can be determined.

[0105] According to an embodiment of the present invention, screening potential research personnel from target personnel based on a first target area and a second target area includes: determining the first centroid coordinate data of the first target area and the second centroid coordinate data of the second target area; determining, based on the calibration parameters of the robot's camera, the first estimated coordinate data of the first centroid coordinate data in a three-dimensional coordinate system with the location of the camera as the origin, and the second estimated coordinate data of the second centroid coordinate data in the three-dimensional coordinate system; and obtaining screening conditions C1 and C2 according to formula (1).

[0106] (1)

[0107] in, For the second estimated coordinate data of the i-th target person, The first estimated coordinate data for the i-th target person. For the preset angle threshold, A preset distance threshold is used; target individuals who meet at least one of the screening conditions C1 and C2 are identified as potential survey participants.

[0108] According to one embodiment of the present invention, since the target person present in both the first and second images may be moving, the yaw angles of the cameras corresponding to the first and second images in which the target person appears may be inconsistent. Therefore, it is difficult to directly determine the movement status of the target person using the image coordinates of the target person in the first and second images. In this case, the first centroid coordinate data of the first target area and the second centroid coordinate data of the second target area can be determined, and the first estimated coordinate data of the first centroid coordinate data in the three-dimensional coordinate system with the location of the camera as the origin, and the second estimated coordinate data of the second centroid coordinate data in the three-dimensional coordinate system can be determined according to the camera calibration parameters (intrinsic and extrinsic parameters). In the example, the height of the target person can be estimated. For example, the height of the target person can be set to a predetermined value (e.g., 1.65 meters). Based on this predetermined value, the first centroid coordinate data, and the camera's intrinsic parameters, the distance between the target person and the camera can be estimated. This allows the determination of the target person's three-dimensional coordinates in the camera coordinate system (a coordinate system with the camera's location as the origin and the direction of the camera's optical axis as the z-axis). Then, the coordinates are adjusted to a unified three-dimensional coordinate system based on the camera's extrinsic parameters (e.g., extrinsic parameters set according to the camera's yaw angle) to obtain the first estimated coordinate data. Similarly, the second estimated coordinate data can be obtained. That is, through the above method, the positions of the same target person in a unified coordinate system at different times can be obtained.

[0109] According to an embodiment of the present invention, in formula (1), Let the vector be the first estimated coordinate data pointing to the second estimated coordinate data, and the direction of this vector be the movement direction of the i-th target person. Let be the vector pointing from the first estimated coordinates to the origin (i.e., the location of the camera), representing the direction of movement required for the i-th target person to approach the camera (robot). Let be the cosine of the angle between the two vectors mentioned above. Therefore, The angle between the two vectors above represents the angle between the movement direction of the i-th target person and the direction of movement required to approach the robot. If the angle is less than or equal to a preset angle threshold (e.g., 45° or 30°), it means that the movement direction of the i-th target person is close to the direction of movement required to approach the robot. That is, the filtering condition C1 can indicate that the i-th target person is moving in a direction closer to the robot; otherwise, it can indicate that the i-th target person is moving in a direction away from the robot.

[0110] According to one embodiment of the present invention, the filtering condition C2 represents the distance the i-th target person moves within the time interval between two times they are photographed. Smaller, that is, less than a preset distance threshold (e.g., 0.5 meters), indicates that the i-th target person is in a state of near stillness.

[0111] According to one embodiment of the present invention, if at least one of the screening conditions C1 and C2 is satisfied, the robot can meet with the i-th target person and conduct a survey of the i-th target person; otherwise, if the i-th target person moves away from the robot at a relatively high speed, it will be difficult for the robot to meet with the i-th target person and conduct a survey of the i-th target person. Therefore, target persons who meet at least one of the screening conditions C1 and C2 can be identified as candidates for survey. Using the same method, it is possible to determine whether each target person is a candidate for survey.

[0112] In this way, the movement of the target personnel can be estimated by estimating their coordinates in a unified three-dimensional coordinate system. Then, screening conditions can be set to select candidates who are moving towards the robot or whose movement speed is slow. This makes it easier for the robot to meet with the candidates and conduct the survey, improving the accuracy of screening candidates and helping to increase the success rate of the survey.

[0113] According to one embodiment of the present invention, in step S4, a first region to be analyzed in the first image where the person to be investigated is located can be determined. For example, the first target region of the person to be investigated can be used as the first region to be analyzed. Similarly, in step S5, the second target region of the person to be investigated can be used as the second region to be analyzed.

[0114] According to one embodiment of the present invention, in step S6, the survey convenience index of the candidate surveyor can be used to describe whether the candidate surveyor is easy to accept the survey. The first and second analysis areas can be analyzed to determine the candidate surveyor's posture. For example, if a candidate surveyor's posture is relatively relaxed, it indicates that the candidate is not busy with other things and is more likely to accept the survey, thus their survey convenience index is high. Conversely, if a candidate surveyor's posture is large (e.g., running quickly) or is handling other things (e.g., reading documents, selecting goods, etc.), their likelihood of accepting the survey is low, and their survey convenience index is low.

[0115] Figure 2 A schematic diagram illustrating the determination of behavior category information according to an embodiment of the present invention is shown as an example.

[0116] According to one embodiment of the present invention, determining the survey convenience index of the candidate survey personnel based on the first region to be analyzed and the second region to be analyzed includes: dividing the first region to be analyzed vertically into a first head region, a first upper body region, and a first lower body region; dividing the second region to be analyzed vertically into a second head region, a second upper body region, and a second lower body region; processing the first head region using a first convolutional neural network model to obtain a first head feature vector, and processing the second head region to obtain a second head feature vector; concatenating the first head feature vector and the second head feature vector to obtain head feature information; inputting the head feature information into a first multilayer perceptron network model to obtain a head feature description vector; processing the first upper body region using a second convolutional neural network model to obtain a first upper body feature vector, and processing the second head region using a second convolutional neural network model to obtain a first upper body feature vector, and ... The upper body region is processed to obtain a second upper body feature vector; the first and second upper body feature vectors are concatenated to obtain upper body feature information; the upper body feature information is input into a second multilayer perceptron model to obtain an upper body feature description vector; the first lower body region is processed through a third convolutional neural network model to obtain a first lower body feature vector, and the second lower body region is processed to obtain a second lower body feature vector; the first and second lower body feature vectors are concatenated to obtain lower body feature information; the lower body feature information is input into a third multilayer perceptron model to obtain a lower body feature description vector; based on the head feature description vector, upper body feature description vector, and lower body feature description vector, the behavioral category information of the candidates to be surveyed is obtained; based on the behavioral category information, the first region to be analyzed, and the second region to be analyzed, the survey convenience index is determined.

[0117] According to one embodiment of the present invention, the actions and postures of the candidate for research can be determined using artificial intelligence, thereby determining whether the candidate is busy with other matters and thus determining the candidate's research convenience index. When determining the actions and postures of the candidate, these actions and postures can be manifested in various parts of their body. For example, the candidate's head, upper body, and lower body can all reflect their ongoing actions and postures. Therefore, accurately identifying the actions and postures of the head, upper body, and lower body helps improve the overall accuracy of identifying the candidate's actions and postures. Therefore, the first region to be analyzed can be divided vertically into a first head region, a first upper body region, and a first lower body region, for example, by dividing it into three equal parts vertically, or by dividing it according to a manually set proportion. Similarly, the second region to be analyzed can be divided into a second head region, a second upper body region, and a second lower body region.

[0118] According to one embodiment of the present invention, a first convolutional neural network model may include multiple convolutional layers, which can process a first head region to obtain a first head feature vector, and process a second head region to obtain a second head feature vector. The first and second head feature vectors are concatenated to obtain head feature information. This head feature information not only describes the head posture of the candidate being surveyed when the first image is taken and the head posture when the second image is taken, but also includes features of the candidate's head movements between the two moments. This allows for a more comprehensive analysis of the candidate's head posture and movements, providing an accurate data foundation for determining the candidate's overall posture and movements and assessing whether the candidate is busy with other tasks. Furthermore, the head feature information can be processed by a first multilayer perceptron model. The head feature information can be a matrix containing two vectors (i.e., a first head feature vector and a second head feature vector). This matrix can be converted into vector form information, i.e., a head feature description vector, by the first multilayer perceptron model. Similarly, an upper body feature description vector can be obtained, which can describe the upper body posture of the candidate researcher when the first image is taken and the upper body posture when the second image is taken, as well as the characteristics of the candidate researcher's upper body movements between the two moments. A lower body feature description vector can also be obtained, which can describe the lower body posture of the candidate researcher when the first image is taken and the lower body posture when the second image is taken, as well as the characteristics of the candidate researcher's lower body movements between the two moments.

[0119] According to one embodiment of the present invention, after obtaining the head feature description vector, upper body feature description vector, and lower body feature description vector, they can be directly concatenated, and the concatenated vector can be input into a fully connected layer and an activation layer (e.g., a layer processed using the softmax activation function) to obtain the behavior category information of the candidate. However, the head movements and postures, upper body movements and postures, and lower body movements and postures are related. The aforementioned head feature description vector, upper body feature description vector, and lower body feature description vector are feature vectors that separately describe the head movements and postures, upper body movements and postures, and lower body movements and postures, respectively, without considering the interrelationships between them. This may cause fragmented recognition results for different body regions, resulting in low accuracy of the behavior category information.

[0120] According to one embodiment of the present invention, in order to represent the correlation between head feature description vectors, upper body feature description vectors, and lower body feature description vectors, a graph neural network model can be used to process these three vectors to describe the correlation between the head, upper body, and lower body. Based on the head feature description vectors, upper body feature description vectors, and lower body feature description vectors, behavioral category information of the candidates to be surveyed is obtained, including: using the head feature description vectors, upper body feature description vectors, and lower body feature description vectors as input vectors for nodes in a graph structure; processing the input vectors of each node through the attention mechanism of the graph neural network model to obtain connection weights between nodes; processing the connection weights and input vectors through the graph neural network model to obtain head feature output vectors, upper body feature output vectors, and lower body feature output vectors; concatenating the head feature output vectors, upper body feature output vectors, and lower body feature output vectors to obtain behavioral description information; and processing the behavioral description information through a fourth multilayer perceptron model to obtain behavioral category information of the candidates to be surveyed.

[0121] According to one embodiment of the present invention, the graph structure may include multiple nodes, and the head, upper body and lower body may be used as nodes of the graph structure respectively. Then the head feature description vector, upper body feature description vector and lower body feature description vector may be used as input vectors of each node respectively.

[0122] According to one embodiment of the present invention, the connection weights between nodes can be used to represent the strength of the association between two nodes. The input vectors of the two nodes can be input into an attention mechanism for processing. The attention mechanism can multiply the input vectors of the two nodes by the weight matrix respectively to obtain two vectors, then concatenate the two vectors, and input the concatenated vector into a fully connected layer and an activation layer (e.g., a layer processed using the sigmoid activation function) to obtain the connection weights between the two nodes.

[0123] According to one embodiment of the present invention, a graph neural network model can use connection weights to weighted sum the input vector of a node and the input vectors of other nodes connected to that node, and then multiply the weighted summed vector by another weight matrix to obtain the output vector of the node. This process is performed on each node to obtain head feature output vectors, upper body feature output vectors, and lower body feature output vectors. Through the above processing using connection weights and weighted summation, the head feature output vector can include not only the head's action and posture features, but also the correlation between the head and the actions and postures of the upper and lower body. Similarly, the upper body feature output vector and the lower body feature output vector can include not only the actions and postures of their respective regions, but also the correlations with other regions.

[0124] According to one embodiment of the present invention, the head feature output vector, upper body feature output vector, and lower body feature output vector can be concatenated to obtain behavioral description information. Then, the behavioral description information is processed by a fourth multilayer perceptron model to obtain behavioral category information of the candidate. The fourth multilayer perceptron model may include multiple levels of fully connected layers and activation layers (e.g., layers processed using the softmax activation function), and can output the probability that the candidate's behavior belongs to multiple categories, and take the category with the highest probability as the candidate's behavioral category information.

[0125] In this way, the first and second regions to be analyzed can be divided, and the common parts after division can be processed to obtain feature description vectors for describing the actions and postures of each part. The correlation between the actions and postures of each part can be obtained through a graph neural network model, thereby improving the accuracy of the overall behavior category information of the researchers to be investigated while accurately identifying the actions and postures of each part.

[0126] According to an embodiment of the present invention, the above-mentioned plurality of neural network models can be trained before use. The method further includes: dividing a first training region in a first training image of a trainee into a first training head region, a first training upper body region, and a first training lower body region; dividing a second training region in a second training image of a trainee into a second training head region, a second training upper body region, and a second training lower body region; obtaining a head training feature vector based on the first training head region and the second training head region; obtaining an upper body training feature vector based on the first training upper body region and the second training upper body region; obtaining a lower body training feature vector based on the first training lower body region and the second training lower body region; obtaining a head training output vector, an upper body training output vector, and a lower body training output vector based on the head training feature vector, the upper body training feature vector, and the lower body training feature vector; inputting the head training output vector into a first fully connected layer and a first activation layer to obtain predicted head action category information; inputting the upper body training output vector into a second fully connected layer and a second activation layer to obtain... The system obtains predicted upper body movement category information; inputs the lower body training output vector into the third fully connected layer and the third activation layer to obtain predicted lower body movement category information; concatenates the head training output vector, upper body training output vector, and lower body training output vector to obtain behavior training information; inputs the behavior training information into the fourth multilayer perceptron model to obtain predicted behavior category information; determines the loss function based on the predicted behavior category information, predicted head movement category information, predicted upper body movement category information, predicted lower body movement category information, and the trainee's annotation information; and trains the first convolutional neural network model, the first multilayer perceptron model, the second convolutional neural network model, the second multilayer perceptron model, the third convolutional neural network model, the third multilayer perceptron model, the graph neural network model, the fourth multilayer perceptron model, the first fully connected layer and the first activation layer, the second fully connected layer and the second activation layer, and the third fully connected layer and the third activation layer using the loss function.

[0127] According to one embodiment of the present invention, during training, an end-to-end training method can be used to train the above-mentioned multiple neural network models as a whole. During training, the acquisition methods of the first training image and the second training image are the same as those for acquiring the first image and the second image. The method of dividing the first training region and the second training region is the same as that for dividing the first region to be analyzed and the second region to be analyzed. The method of acquiring the head training feature vector, upper body training feature vector, and lower body training feature vector is the same as that for acquiring the head feature description vector, upper body feature description vector, and lower body feature description vector. The method of obtaining the head training output vector, upper body training output vector, and lower body training output vector is the same as that for obtaining the head feature output vector, upper body feature output vector, and lower body feature output vector, and will not be elaborated further here.

[0128] According to one embodiment of the present invention, the head training output vector can not only describe the head's movements and postures, but also represent the correlation between the head and the movements and postures of the upper and lower body, thus more accurately describing the head's work categories. Therefore, the head training output vector can be input into a first fully connected layer and a first activation layer (e.g., a layer processed using the softmax activation function) to output predicted head movement category information, that is, the probability that the head's movements belong to multiple movement categories (e.g., the probability that the head is looking down while reading, staring at something, or looking around). Similarly, the upper body training output vector can be input into a second fully connected layer and a second activation layer to obtain predicted upper body movement category information, that is, the probability that the upper body's movements belong to multiple movement categories (e.g., the probability of swinging arms while walking, waving arms while running, carrying objects, or holding a mobile phone). The lower body training output vector can also be input into a third fully connected layer and a third activation layer to obtain predicted lower body movement category information (e.g., the probability of leg movements while walking, leg movements while running, standing, or sitting).

[0129] According to one embodiment of the present invention, the head training output vector, the upper body training output vector, and the lower body training output vector can be concatenated to obtain behavioral training information, and then input into a fourth multilayer perceptual network model to obtain predicted behavior category information, that is, the probability that the trainee's behavior belongs to various behavior categories (e.g., the probability of looking around, talking, running, reading, etc.).

[0130] According to one embodiment of the present invention, a loss function can be determined based on the various probability information determined above, so as to improve the prediction accuracy of actions of various parts and the prediction accuracy of overall behavior. The loss function is determined based on the predicted behavior category information, predicted head action category information, predicted upper body action category information, predicted lower body action category information, and the annotation information of the trainee, including: determining the loss function LOSS according to formula (2).

[0131] (2)

[0132] in, Let be the probability that the trainee's head movement, determined based on the annotation information, is the k-th type of movement. Let be the probability that a trainee's head movement is the k-th type, determined based on the predicted head movement category information. Let be the probability that the trainee's upper body movement, determined based on the annotation information, is the t-th type of movement. Let be the probability that a trainee's upper body movement is the t-th type, determined based on the predicted upper body movement category information. Let s be the probability that the trainee's lower body movement is the s-th type, determined based on the annotation information. Let s be the probability that a trainee's lower body movement is the s-th type, determined based on the predicted lower body movement category information. Let $\mathbf{j}$ be the probability that the behavior category of the trainee, determined based on the annotation information, belongs to the $j$ category. Let $\mathbf{j}$ be the probability that the trainee's behavior category is the $j$-th category, determined based on the predicted behavior category information. The number of categories of head movements. The number of categories of upper body movements. The number of categories of lower body movements. For the number of behavior categories, , , and For the preset weights, k≤ , t≤ s≤ j≤ And k, ,t, ,s, j All are positive integers.

[0133] According to one embodiment of the present invention, when the trainee's head movement is the kth type of movement, ,otherwise, .therefore, for and The constructed cross-entropy loss function decreases during training, resulting in improved model prediction accuracy. Closer Furthermore, by summing the cross-entropy loss functions for various head movement types, we can obtain the head movement loss function. .

[0134] According to one embodiment of the present invention, when the trainee's upper body movement is the t-th movement, ,otherwise, .therefore, for and The constructed cross-entropy loss function decreases during training, resulting in improved model prediction accuracy. Closer Furthermore, by summing the cross-entropy loss functions for various upper body movement types, we can obtain the upper body movement loss function. .

[0135] According to one embodiment of the present invention, when the trainee's lower body movement is the s-th type of movement, ,otherwise, .therefore, for and The constructed cross-entropy loss function decreases during training, resulting in improved model prediction accuracy. Closer Furthermore, by summing the cross-entropy loss functions for various lower body movement types, we can obtain the lower body movement loss function. .

[0136] According to one embodiment of the present invention, when the trainer's behavior category is the j-th category, ,otherwise, .therefore, for and The constructed cross-entropy loss function decreases during training, resulting in improved model prediction accuracy. Closer Furthermore, by summing the cross-entropy loss functions for multiple behavior categories, we can obtain the behavior category loss function. .

[0137] According to one embodiment of the present invention, a weighted sum of the head action loss function, upper body action loss function, lower body action loss function, and behavior category loss function can be obtained to obtain the final loss function. This loss function can be used for backpropagation to update the parameters of the first convolutional neural network model, the first multilayer perceptron model, the second convolutional neural network model, the second multilayer perceptron model, the third convolutional neural network model, the third multilayer perceptron model, the graph neural network model, the fourth multilayer perceptron model, the first fully connected layer and the first activation layer, the second fully connected layer and the second activation layer, and the third fully connected layer and the third activation layer, through gradient descent, in order to train these neural network models. After multiple training iterations, the training can be completed, and the above neural network models can be used to predict the behavior category information of the survey participants.

[0138] In this way, the head training output vector, upper body training output vector, and lower body training output vector can be optimized separately. By combining the correlation between various parts, the accuracy of motion prediction for each part can be improved, thereby increasing the accuracy of the head, upper body, and lower body training output vectors, enhancing training intensity, and providing an accurate data foundation for obtaining more accurate predicted behavior category information. Furthermore, the predicted behavior category information can be optimized as a whole, thereby training multiple neural network models as a whole, enhancing training intensity, and improving the prediction accuracy of behavior category information.

[0139] According to one embodiment of the present invention, after obtaining accurate behavioral category information, a survey convenience index for a candidate surveyor can be determined. Determining the survey convenience index based on the behavioral category information, the first region to be analyzed, and the second region to be analyzed includes: determining, based on the first centroid coordinate data of the first region to be analyzed, the first estimated coordinate data of the candidate surveyor in a three-dimensional coordinate system with the camera's location as the origin, and the second estimated coordinate data of the second centroid coordinates of the second region to be analyzed in the same three-dimensional coordinate system; determining, based on the first and second estimated coordinate data, the predicted motion trajectory data and motion speed data of the candidate surveyor in the three-dimensional coordinate system; and determining the survey convenience index based on the behavioral category information, the predicted motion trajectory data, the motion speed data, and the robot's motion speed.

[0140] According to one embodiment of the present invention, the method for obtaining the first estimated coordinate data and the second estimated coordinate data is the same as described above, and will not be repeated here. The predicted trajectory data of the candidate in the three-dimensional coordinate system can be set as a trajectory of uniform linear motion along the direction of the vector pointing from the first estimated coordinate data to the second estimated coordinate data. The motion speed data is equal to the magnitude of the vector pointing from the first estimated coordinate data to the second estimated coordinate data, and the ratio of the time difference between the moment the candidate is captured in the first round of panoramic view and the moment the candidate is captured in the second round of panoramic view. If the candidate is captured more than once in each round of panoramic view, the moment the candidate is captured in the image closest to the center of the image is taken as the moment the candidate is captured in that round of panoramic view. For example, in the first round of panoramic view, the candidate is captured in both the second and third first images. In order to reduce the interference of distortion caused by image edges, the first image closest to the center of the image is used as the first image of the candidate being captured, and the moment the first image is captured is taken as the moment the candidate is captured in the first round of panoramic view. The first estimated coordinate data is the coordinate of the candidate at this moment. Similarly, the determination method of the second estimated coordinate and the moment the candidate is captured in the second round of panoramic view can be obtained.

[0141] According to one embodiment of the present invention, determining a survey convenience index based on the behavior category information, the predicted motion trajectory data, the motion speed data, and the robot's motion speed includes: determining the survey convenience index of the g-th candidate surveyor according to formulas (3), (4), (5), and (6). ,

[0142] (3)

[0143] (4)

[0144] (5)

[0145] (6)

[0146] in, Let be the probability that the behavior category determined based on the behavior category information of the g-th undetermined survey participant is the j-th category. For the number of behavior categories, The preset convenience coefficient for the behavior of the j-th category. The time required for the robot to meet with the g-th undetermined researcher. For the g-th undetermined survey participant, the second estimated coordinate data represents the planar coordinates. For the g-th undetermined survey personnel, the first estimated coordinate data represents the planar coordinates. Let g be the angle of motion of the g-th undetermined researcher. For the movement speed data of the g-th undetermined researcher, , The time difference between the moment when the undetermined research personnel were captured during the first round of the surround view and the moment when they were captured during the second round of the surround view. For the robot's movement speed, Let j be the robot's direction angle, j≤ And j and All are positive integers.

[0147] According to one embodiment of the present invention, the predicted motion trajectory data of the candidate researcher in the three-dimensional coordinate system can be set as a trajectory of uniform linear motion along the direction of the vector pointing from the first estimated coordinate data to the second estimated coordinate data. Therefore, when the robot meets the g-th candidate researcher, the planar coordinates of the position of the g-th candidate researcher are: The robot's coordinates are Since the two converge, the two coordinates can be made equal, that is, ,therefore, , After sorting, we can obtain... ,therefore, ,because ,therefore, ,Right now, Formula (5) can then be obtained. Equation (4) can be obtained by simplification. Substituting equations (5) and (6) into equation (4), the time required for the robot to meet with the g-th undetermined survey participant can be obtained. Therefore, the survey convenience index can be solved by substituting it into formula (3).

[0148] According to one embodiment of the present invention, in formula (3), the preset convenience coefficient can be a weight set manually for each behavior category. The more convenient the candidate is to accept the survey when performing the behavior, the higher the preset convenience coefficient for that behavior category. For example, when the behavior category is "looking around," the preset convenience coefficient is 0.3; when the behavior category is "waiting still," the preset convenience coefficient is 0.4; when the behavior category is "reading," the preset convenience coefficient is 0.05; when the behavior category is "walking," the preset convenience coefficient is 0.2; and when the behavior category is "running," the preset convenience coefficient is 0.05. By using the preset convenience coefficient and the probability of the g-th candidate performing various behaviors to perform weighted summation, the index of the g-th candidate's ease of accepting the survey can be obtained. This index is related to... The ratio can be used as an indicator of survey convenience. For example, if a candidate has a high indicator of being easy to survey, and is relatively close to the robot, then... If the size of the data is relatively small, it is easy for the robot to move to the vicinity of the person to be surveyed, and the person to be surveyed is convenient to accept the survey, then the survey convenience index of the person to be surveyed is relatively high.

[0149] In this way, the predicted motion trajectory and speed data of the candidates to be surveyed can be set, and the time required for the candidates to meet with the robot can be calculated. Then, the survey convenience index can be determined by the behavior category information and the time required for meeting. This allows for the selection of candidates who are easy to meet with the robot and easy to accept the survey, thereby improving the efficiency and quality of the survey.

[0150] According to an embodiment of the present invention, in step S7, target survey personnel can be screened based on survey convenience indicators. For example, the candidate with the highest survey convenience indicator can be determined as the target survey personnel. In step S8, the robot's movement component is controlled to move the robot to a preset distance in front of the target survey personnel's face and issue a survey invitation message, for example, by issuing a voice invitation message to invite the target survey personnel to participate in the survey. If the target survey personnel accept the survey, the survey questionnaire can be displayed on the display device. If the target survey personnel refuse, or if the target survey personnel leave the robot's camera's field of vision during the robot's movement, making the robot unable to detect the target survey personnel, the above process can be repeated to find a new target survey personnel.

[0151] According to embodiments of the present invention, a vision-based robot-based automated survey method can screen potential survey participants using first and second images captured by the robot's camera during two rounds of surround view. It automatically analyzes whether a participant is easily surveyed based on their location, thereby reducing the randomness of participant selection and automatically choosing those most likely to be surveyed, thus improving survey quality. When screening potential participants, the method estimates their movement by estimating their coordinates in a unified three-dimensional coordinate system, and then sets screening criteria to identify participants moving towards the robot or moving at a slower speed. This facilitates the robot's convergence with the participants and improves the accuracy of participant selection, increasing the survey success rate. When determining behavioral category information, the method divides the first and second analysis areas and processes identical parts to obtain feature vectors describing the actions and postures of each part. A graph neural network model can then be used to obtain the correlation between the actions and postures of each part, thereby improving the accuracy of the overall behavioral category information of the potential participants while accurately identifying the actions and postures of each part. When training neural network models, the head, upper body, and lower body training output vectors can be optimized separately. By combining the correlations between these parts, the accuracy of motion prediction for each part can be improved, thereby enhancing the accuracy of the head, upper body, and lower body training output vectors. This increases training intensity and provides a more accurate data foundation for obtaining more precise predicted behavior category information. Furthermore, the predicted behavior category information can be optimized holistically to train multiple neural network models as a whole, further increasing training intensity and improving the accuracy of behavior category prediction. When determining survey convenience indicators, the predicted motion trajectory and speed data of the candidates can be set, and the time required for the candidates to meet with the robot can be calculated. Then, the survey convenience indicators can be determined by using behavior category information and the required meeting time. This allows for the selection of candidates who are easy to meet with the robot and convenient to accept the survey, improving the survey's efficiency and quality.

[0152] Figure 3 An exemplary block diagram of a vision-based automated robotic survey system according to an embodiment of the present invention is shown, the system comprising:

[0153] The first image module is used to capture first images from multiple angles during the first round of surround view using the robot's camera;

[0154] The second image module is used to capture second images from multiple angles during the second round of surround view using the robot's camera;

[0155] A screening module is used to select potential survey participants from multiple people around the robot based on the first image and the second image.

[0156] The first region to be analyzed module is used to obtain the first region to be analyzed in the first image where the person to be surveyed is located.

[0157] The second region to be analyzed module is used to obtain the second region to be analyzed in the second image where the person to be surveyed is located.

[0158] The survey convenience index module is used to determine the survey convenience index of the candidates based on the first region to be analyzed and the second region to be analyzed.

[0159] The target survey personnel module is used to select target survey personnel from the undetermined survey personnel based on the survey convenience index.

[0160] The survey invitation module is used to control the robot's movement components, causing the robot to move to a preset distance in front of the target survey participant's face and send a survey invitation message.

[0161] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0162] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been shown and explained in the embodiments, and any modifications or variations of the implementation of the present invention may be made without departing from the stated principles.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual inspection-based automated surveying method for robots, characterized in that, include: The robot's camera captures first images from multiple angles during the initial surround view process; The robot's camera captures second images from multiple angles during the second round of surround view. Based on the first image and the second image, select potential research subjects from among multiple people around the robot; Obtain the first region to be analyzed in the first image where the researchers to be identified are located. Obtain the second region to be analyzed in the second image where the researchers to be identified are located; Based on the first region to be analyzed and the second region to be analyzed, determine the survey convenience index for the candidates to be surveyed; Based on the aforementioned survey convenience index, target survey personnel are selected from the pool of undetermined survey personnel; Control the robot's movement components to move the robot to a preset distance in front of the target researcher's face and send out a research invitation message; Based on the first and second regions to be analyzed, determine the survey convenience indicators for the candidates to be surveyed, including: The first region to be analyzed is divided into a first head region, a first upper body region, and a first lower body region in the vertical direction; The second region to be analyzed is divided vertically into a second head region, a second upper body region, and a second lower body region. The first head region is processed using the first convolutional neural network model to obtain the first head feature vector, and the second head region is processed to obtain the second head feature vector. The first head feature vector and the second head feature vector are concatenated to obtain head feature information; The head feature information is input into the first multilayer perceptron model to obtain the head feature description vector. The first upper body region is processed using a second convolutional neural network model to obtain the first upper body feature vector, and the second upper body region is processed to obtain the second upper body feature vector. The first upper body feature vector and the second upper body feature vector are concatenated to obtain upper body feature information; The upper body feature information is input into the second multilayer perceptron model to obtain the upper body feature description vector; The first lower body region is processed using a third convolutional neural network model to obtain the first lower body feature vector, and the second lower body region is processed to obtain the second lower body feature vector. The first lower body feature vector and the second lower body feature vector are concatenated to obtain lower body feature information; The lower body feature information is input into the third multilayer perceptron model to obtain the lower body feature description vector; Based on the head feature description vector, upper body feature description vector, and lower body feature description vector, obtain the behavioral category information of the researchers to be determined; Based on the behavioral category information, the first region to be analyzed, and the second region to be analyzed, the survey convenience index is determined.

2. The robot-based automated survey method according to claim 1, characterized in that, Based on the first image and the second image, select potential research subjects from among multiple people around the robot, including: Detection is performed on the first image from multiple angles to obtain the first region where the person is located in each first image, wherein the first region is a rectangular area that selects the range of the person's location; By using an image detection model, feature extraction processing is performed on the first region to obtain the first appearance feature information of each person; Detection is performed on the second images from multiple angles to obtain the second region where the person is located in each second image; The second region is processed by an image detection model to obtain the second appearance feature information of each person. Based on the first and second facial feature information, a target person is identified that exists in both the first and second images. Obtain the first target region of the target person in the first image and the second target region in the second image; Based on the first and second target areas, select potential research personnel from among the target personnel.

3. The robot-based automated survey method according to claim 2, characterized in that, Based on the first and second target areas, select potential research participants from the target population, including: Determine the first centroid coordinate data of the first target region and the second centroid coordinate data of the second target region; Based on the calibration parameters of the robot's camera, determine the first estimated coordinate data of the first centroid coordinate data in a three-dimensional coordinate system with the location of the camera as the origin, and the second estimated coordinate data of the second centroid coordinate data in the three-dimensional coordinate system. According to the formula Obtain the filtering conditions C1 and C2, where, For the second estimated coordinate data of the i-th target person, The first estimated coordinate data for the i-th target person. For the preset angle threshold, The preset distance threshold; Individuals who meet at least one of the screening criteria C1 and C2 are identified as potential survey participants.

4. The robot-based automated survey method according to claim 1, characterized in that, Based on the head feature description vector, upper body feature description vector, and lower body feature description vector, behavioral category information of the individuals to be surveyed is obtained, including: The head feature description vector, upper body feature description vector, and lower body feature description vector are used as the input vectors of the nodes in the graph structure, respectively. By using the attention mechanism of the graph neural network model, the input vectors of each node are processed to obtain the connection weights between nodes. The connection weights and input vectors are processed by a graph neural network model to obtain head feature output vectors, upper body feature output vectors, and lower body feature output vectors. The head feature output vector, upper body feature output vector, and lower body feature output vector are concatenated to obtain behavioral description information; The behavioral description information is processed by the fourth multilayer perceptual network model to obtain the behavioral category information of the researchers to be surveyed.

5. The visual inspection-based automated robot survey method according to claim 4, characterized in that, The method further includes: The first training region of the trainee in the first training image is divided into the first training head region, the first training upper body region, and the first training lower body region. The second training region in the second training image is divided into the second training head region, the second training upper body region, and the second training lower body region. Based on the first training head region and the second training head region, obtain the head training feature vector; based on the first training upper body region and the second training upper body region, obtain the upper body training feature vector; based on the first training lower body region and the second training lower body region, obtain the lower body training feature vector. Based on the head training feature vector, upper body training feature vector, and lower body training feature vector, obtain the head training output vector, upper body training output vector, and lower body training output vector; The head training output vector is input into the first fully connected layer and the first activation layer to obtain the predicted head action category information; The upper body training output vector is input into the second fully connected layer and the second activation layer to obtain the predicted upper body movement category information; The lower body training output vector is input into the third fully connected layer and the third activation layer to obtain the predicted lower body movement category information; The head training output vector, upper body training output vector, and lower body training output vector are concatenated to obtain behavioral training information. The behavioral training information is input into the fourth multilayer perceptron model to obtain the predicted behavioral category information. The loss function is determined based on the predicted behavior category information, predicted head movement category information, predicted upper body movement category information, predicted lower body movement category information, and the annotation information of the trainees. The loss function is used to train a first convolutional neural network model, a first multilayer perceptron model, a second convolutional neural network model, a second multilayer perceptron model, a third convolutional neural network model, a third multilayer perceptron model, a graph neural network model, a fourth multilayer perceptron model, a first fully connected layer and a first activation layer, a second fully connected layer and a second activation layer, and a third fully connected layer and a third activation layer.

6. The robot-based automated survey method according to claim 5, characterized in that, Based on the predicted behavior category information, predicted head movement category information, predicted upper body movement category information, predicted lower body movement category information, and the trainer's annotation information, the loss function is determined, including: According to the formula Determine the loss function LOSS, where, Let be the probability that the trainee's head movement, determined based on the annotation information, is the k-th type of movement. Let be the probability that a trainee's head movement is the k-th type, determined based on the predicted head movement category information. Let be the probability that the trainee's upper body movement, determined based on the annotation information, is the t-th type of movement. Let be the probability that a trainee's upper body movement is the t-th type, determined based on the predicted upper body movement category information. Let s be the probability that the trainee's lower body movement is the s-th type, determined based on the annotation information. Let s be the probability that a trainee's lower body movement is the s-th type, determined based on the predicted lower body movement category information. Let $\mathbf{j}$ be the probability that the behavior category of the trainee, determined based on the annotation information, belongs to the $j$ category. Let $\mathbf{j}$ be the probability that the trainee's behavior category is the $j$-th category, determined based on the predicted behavior category information. The number of categories of head movements. The number of categories of upper body movements. The number of categories of lower body movements. For the number of behavior categories, , , and For the preset weights, k≤ , t≤ s≤ j≤ And k, ,t, ,s, j All are positive integers.

7. The robot-based automated survey method according to claim 1, characterized in that, Based on the behavioral category information, the first region to be analyzed, and the second region to be analyzed, survey convenience indicators are determined, including: Based on the first centroid coordinate data of the first region to be analyzed, the first estimated coordinate data of the personnel to be investigated in the three-dimensional coordinate system with the location of the camera as the origin is determined, and the second estimated coordinate data of the second centroid coordinates of the second region to be analyzed in the three-dimensional coordinate system is determined. Based on the first and second estimated coordinate data, the predicted motion trajectory and motion speed data of the personnel to be investigated in the three-dimensional coordinate system are determined. Based on the behavior category information, the predicted motion trajectory data, the motion speed data, and the robot's motion speed, a survey convenience index is determined.

8. The visual inspection-based automated robot survey method according to claim 7, characterized in that, Based on the behavior category information, the predicted motion trajectory data, the motion speed data, and the robot's motion speed, survey convenience indicators are determined, including: According to the formula Determine the survey convenience index for the g-th undetermined survey participant. ,in, Let be the probability that the behavior category determined based on the behavior category information of the g-th undetermined survey participant is the j-th category. For the number of behavior categories, The preset convenience coefficient for the behavior of the j-th category. The time required for the robot to meet with the g-th undetermined researcher. For the g-th undetermined survey participant, the second estimated coordinate data represents the planar coordinates. For the g-th undetermined survey personnel, the first estimated coordinate data represents the planar coordinates. Let g be the angle of motion of the g-th undetermined researcher. For the movement speed data of the g-th undetermined researcher, , The time difference between the moment when the undetermined research personnel were captured during the first round of the surround view and the moment when they were captured during the second round of the surround view. For the robot's movement speed, Let j be the robot's direction angle, j≤ And j and All are positive integers.

Citation Information

Patent Citations

  • Robot management system

    CN109129460A

  • Automatic conference investigation robot

    CN119550363A