A method for identifying the content of hair-cutting services based on activity diagrams
The barber service content recognition method is constructed through graph convolution neural network technology, combined with face and tool detection, and a behavioral activity diagram is established, which solves the problems of false detection and missed detection of barber service content recognition in the existing technology, and realizes intelligent management of barber shops.
Patent Information
- Application Number
- CN202211467539.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-22
AI Technical Summary
The existing hairdressing service content recognition methods have a high probability of false detection and missed detection in surveillance videos. It is impossible to effectively combine the relationship between barbers, customers and hairdressing tools and other elements, making it difficult to achieve intelligent management of barbershops.
Using graph convolution neural network technology, the barber’s face image library, the hairdresser tool image library and the standard hairdressing behavior activity library are established, and the customer’s hairdressing behavior activity diagram is constructed by combining face recognition and object detection models, and the graph similarity calculation model is used to identify the content of the hairdressing service.
It realizes the rapid and accurate identification of the content of hairdressing services, improves the intelligent management efficiency of barber shops, and reduces the probability of false inspection and missed inspection.
Smart Images

Figure CN116229342B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and relates to a method for identifying the content of hair-cutting services based on a behavior activity graph. Background Art
[0002] With the increasing maturity of large hair-cutting chain stores, the business management and service management in hair salons are gradually tending towards informatization and intelligence to improve work efficiency and management effectiveness. Under the guidance of the demand for intelligent services and management, it has become an important link for large hair-cutting chain stores to intelligently supervise the authenticity of the work content and service attitude of staff. Among them, the authenticity of the service work content in the hair salon needs to be achieved through video image processing and recognition to achieve the supervision purpose. The recognition of hair-cutting service content based on video images has become an important technical means.
[0003] In order to reduce the probability of false detection and missed detection in the recognition of hair-cutting service content, it is far from enough to only use the existing behavior recognition model to recognize the service behavior of the barber in the surveillance video. It is necessary to establish a behavior activity graph and its recognition model by combining the mutual relationships among elements such as barbers, customers, hair-cutting tools, and the behavior activity process to assist in realizing the remote real-time intelligent monitoring of the hair salon and achieving the purpose of intelligent management. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for identifying the content of hair-cutting services based on a behavior activity graph, which mainly adopts graph convolutional neural network technology and can accurately identify the hair-cutting service content enjoyed by customers in the hair salon. The similarity search of graphs is one of the most important graph-based applications. The similarity calculation of graphs, such as graph edit distance (GED) and maximum common subgraph, is the core operation of graph similarity search and many other applications, but the computational complexity is very high in practice. In recent years, graph neural networks (Graph Neural Networks, GNNs) have shown excellent capabilities in capturing topological information in data and have achieved excellent results in the fields of computer vision and natural language processing, so they have received extensive attention from researchers. Graph convolutional neural networks (GCNs) are a class of graph neural networks that perform outstandingly and are widely applied in GNNs, and are very suitable for processing irregular graph-structured data.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A method for identifying the content of hair-cutting services based on a behavior activity graph, comprising the following steps:
[0007] S1: Set the relative positions of the hair-cutting seats and the video monitoring devices; establish a barber face image library, a hair-cutting tool image library, and a standard hair-cutting behavior activity graph library for detecting and identifying the barber's identity, hair-cutting tools, and hair-cutting service content;
[0008] S2: Confirm the identities of the barber and the customer through the face recognition model and the object tracking model; Detect and identify the haircut tools that appear in the video frame through the object detection model. Mark their service behaviors according to the corresponding relationship with the service behaviors.
[0009] S3: In each frame, calculate the Euclidean distance between the customer detection box and the haircut tool detection box. According to the corresponding relationship between the haircut tool and the service behavior, mark the service behavior of each frame with the haircut tool detection box when the distance is the smallest. Through the above processing of the video information of the service process, construct the customer haircut behavior activity diagram;
[0010] S4: Construct and train a haircut service content recognition model based on graph similarity, and match the customer haircut behavior activity diagram with the standard haircut behavior activity diagram to determine the service content received by the customer in the barbershop.
[0011] Furthermore, in the step S1, the following specific steps are included:
[0012] S11: Set the relative position between the haircut seat and the monitoring camera, and specify the hardware conditions of the monitoring camera to meet the requirements of model accuracy;
[0013] S12: Establish a barber face image library, where each barber face image is named after the corresponding barber identity name for identifying the barber face in the video frame; Establish a haircut tool image library, where each haircut tool image is named after the corresponding service behavior name for identifying the service behavior in the video frame, and the service behavior and the haircut tool are in a one-to-many relationship;
[0014] S13: Establish a standard haircut behavior activity diagram library: A haircut service content consists of several service behaviors. The standard haircut behavior activity diagram constructed by these service behaviors is represented as: G barc_i =<V barc_i , E barc_i >, where G barc_i is an undirected graph corresponding to the barc_i-th haircut service content, barc_i ∈ [0, barc_n - 1], and barc_n is the number of types of haircut service content; V barc_i is the service behavior node set of G barc_i : V barc_i ={v barc_i_x |barc_i_x ∈ [0, barc_i_n - 1]}, barc_i_n is the number of elements in V barc_i , and each element in V barc_i stores a service behavior name; The edge set of G barc_i : E barc_i ={(v barc_i_x, v barc_i_y ) | barc_i_x, barc_i_y ∈ [0, barc_i_n - 1]}, E barc_i Each element in represents v barc_i_x The corresponding service behavior occurs prior to v barc_i_y The corresponding service behavior.
[0015] Furthermore, the steps for establishing the standard haircut behavior activity picture library are as follows:
[0016] S131: According to the standard service process description corresponding to each haircut service content in the barbershop, shoot the service demonstration video for each haircut service content under the same conditions;
[0017] S132: Mark the service behavior of each frame in the demonstration video corresponding to the barc_i-th haircut service content by time. According to the marked service behavior types, add the service behavior nodes to the service behavior node set V barc_i of G barc_i , and at the same time, for the service behavior node v barc_i_x corresponding to a certain frame and the service behavior node v barc_i_y corresponding to the next frame, add (v barc_i_x , v barc_i_y ) to the edge set E barc_i of G barc_i . Finally, form the complete standard haircut behavior activity graph G barc_i for the barc_i-th haircut service content;
[0018] S133: Use the method described in step S132 to establish the corresponding standard haircut behavior activity graphs for all haircut service contents, and form the standard haircut behavior activity picture library.
[0019] Furthermore, in step S2, the following steps are included:
[0020] S21: Set a certain range where the face detection frame appears in the video frame after the customer sits down as the fixed detection frame range for service content video detection. The fixed detection frame is represented as: Frame det = (x det , y det , w det , h det ), where x det and y det respectively represent the horizontal and vertical axis coordinates of the upper left corner point of the fixed detection frame, and w det and h det respectively represent the number of pixels of the fixed detection frame along the vertical and horizontal axes, that is, the width and length of the fixed detection frame;
[0021] S22: Use the face recognition model and the target tracking model to track the faces in the surveillance video. The face recognition model is used to detect the faces in the video frames and identify the barber faces among them. The target tracking model is used to associate the face bounding boxes output by the face recognition model in each video frame to achieve face tracking. The specific implementation steps of face tracking are as follows: Use the barber face image library to perform offline pre-training on the face recognition model, so that the face recognition model can extract the positions and features of the faces in each frame. At the same time, assign the successfully recognized barber faces to the corresponding barber identity names, assign non-repetitive customer IDs to the unrecognized faces, and save the features of all current faces; Use the target tracking model to first predict the positions of the faces in the current frame, and update the positions of the face bounding boxes in the current frame through the similarity calculation module and the data association module of the target tracking model for the prediction results and the face detection results of the face recognition model. Calculate and match the face features in the current frame with the face features saved in the previous face tracking, and assign the same identity information to each matched face to achieve face tracking;
[0022] S23: Confirm the customer detection bounding boxes within the fixed detection bounding box range generated in step S22 for each frame. The specific calculation steps are as follows: Let the video frame in which the face is first detected be the first frame. Then, the ci-th customer detection bounding box in the t-th frame is expressed as: Frame cus_t_ci =(x cus_t_ci ,y cus_t_ci ,w cus_t_ci ,h cus_t_ci ,id cus_t_ci ), where ci ∈ [0, cus_t_cn - 1], cus_t_cn is the total number of customer detection bounding boxes in the t-th frame obtained in step S22, x cus_t_ci and y cus_t_ci represent the horizontal and vertical coordinates of the upper left corner point of Frame cus_t_ci , w cus_t_ci and h cus_t_ci represent the width and length of Frame cus_t_ci ; id cus_t_ci represents the customer ID of Frame cus_t_ci . Therefore, the center coordinates of Frame cus_t_ci can be obtained as If and Then confirm that the current customer is the customer receiving the haircut service, and save the customer ID of the customer detection bounding box confirmed in the t-th frame and the customer detection bounding box index of this box: id cus_t = id cus_t_ci , index cus_t = cus_t_ci;
[0023] S24: Use the haircut tool image library to perform offline pre-training on the object detection model so that the object detection model can output the positions of haircut tools in each video frame and label them with corresponding service behavior names. Similar to step S23, assume that the video frame in which the face is first detected is the first frame. Then, the bi-th haircut tool detection box in the t-th frame is represented as: Frame bart_t_bi =(x bart_t_bi ,y bart_t_bi ,w bart_t_bi ,h bart_t_bi ,p bart_t_bi ), where bi ∈ [0, bart_t_bn - 1], and bart_t_bn is the total number of haircut tool detection boxes in the t-th frame; x bart_t_bi and y bart_t_bi represent the horizontal and vertical axis coordinates of the upper left corner point of Frame bart_t_bi , w bart_t_bi and h bart_t_bi represent the width and length of Frame bart_t_bi , and p bart_t_bi represents the service behavior name corresponding to Frame bart_t_bi .
[0024] Furthermore, in step S3, it specifically includes the following steps:
[0025] S31: Similar to step S23, assume that the video frame in which the face is first detected is the first frame; assume that the identity ID of the first confirmed customer after the first frame is id cus_now . Starting from the frame in which the customer's identity ID is first confirmed as id cus_now , compare the id cus_t saved in each frame with id cus_now until the frame where id cus_t is empty for a certain number of consecutive frames or the frame where the non-empty id cus_t is not equal to id cus_now . Calculate the Euclidean distance between the center of each haircut tool detection box in each frame and the center of the customer detection box with the identity ID of id cus_now . Finally, obtain the service behavior name of the haircut tool detection box participating in the calculation when the Euclidean distance is the smallest, and identify it as the service behavior associated with each frame of the customer with the identity ID of id cus_now .
[0026] S32: According to the service behaviors identified in each frame with the confirmed identity ID of id cus_now , establish a haircut behavior activity diagram for the customer with the identity ID of id cus_now as the input of the haircut service content recognition model; the haircut behavior activity diagram for the customer with the identity ID of id cus_now is represented as: Among them is an undirected graph, is the set of service behavior nodes of: id cus_now _n is the number of elements in, each element in stores a service behavior name; E' idcus_now is the edge set of: each element in represents the corresponding service behavior occurs prior to the corresponding service behavior.
[0027] Furthermore, the hair-cutting behavior association steps in step S31 are as follows:
[0028] S311: The center coordinates of the hair-cutting tool detection box Frame bart_t_bi in the t-th frame are The customer detection box with the confirmed identity ID as id cus_now in the t-th frame is represented as:
[0029]
[0030] where index cus_now_t represents the index of the customer detection box with the identity ID as id cus_now in the customer detection boxes in the t-th frame; and represent the horizontal and vertical axis coordinates of the upper left corner point; and represent the width and length of; The center coordinates of are Furthermore, the Euclidean distance between the center of Frame bart_t_bi in the t-th frame and the center of can be obtained as
[0031]
[0032] S312: Let the minimum Euclidean distance calculated in the t-th frame be d t_min , and the index of the hair-cutting tool detection box participating in the calculation corresponding to it is index bart_t ; Calculate the Euclidean distance between the center of each hair-cutting tool detection box in the t-th frame and the center of in accordance with the calculation method provided in step S311 until all hair-cutting tool detection boxes in the t-th frame are traversed; During the loop, each calculated Euclidean distance is compared with the current value of d t_min ; if it is less than dt_min , then d t_min is updated to the currently calculated Euclidean distance, and at the same time index bart_t is also updated to the index of the haircut tool detection box of the haircut tool participating in the calculation in the t-th frame; when the loop calculation ends, save the service behavior name of the index bart_t th haircut tool detection box in the t-th frame: , and complete the recognition of the service behavior associated with the customer with the identity ID of id cus_now in the t-th frame.
[0033] Furthermore, the steps for establishing the actual haircut behavior activity graph in step S32 include:
[0034] Starting from the frame where the customer identity ID is first confirmed as id cus_now , compare the id cus_t saved in each frame with id cus_now until the frame where id cus_t is empty for a certain number of consecutive frames or the frame where the non-empty id cus_t is not equal to id cus_now ends. Traverse the service behavior name p t_min recognized in each frame, and according to the type of recognized service behavior, add the service behavior node to the service behavior node set V' idcus_now . At the same time, for the service behavior node corresponding to a certain frame and the behavior node corresponding to the next frame, add to the edge set . Finally, form the haircut behavior activity graph of the customer with the customer identity ID of id cus_now . .
[0035] Furthermore, in step S4, the following steps are specifically included:
[0036] S41: Construct a training dataset for the haircut behavior activity graph for offline pre-training of the haircut service content recognition model; the specific construction process of the dataset is as follows: Obtain multiple customer haircut behavior activity graphs in a real haircut scenario through steps S2 - S3, combine each obtained customer haircut behavior activity graph with each standard haircut behavior activity graph in the standard haircut behavior activity graph library to form multiple graph pairs, perform graph edit distance (GED) calculation on the customer haircut behavior activity graph and the standard haircut behavior activity graph in each graph pair using the A* algorithm, and finally combine the node set and edge set of each graph in each graph pair and the GED of the two graphs to form training data. All the training data constitutes the training dataset for the haircut behavior activity graph;
[0037] S42: Construct a hair styling service content recognition model for recognizing the hair styling service content that a customer enjoys in a barbershop; the hair styling service content recognition model includes a GOTSim graph similarity calculation model and a hair styling service content output module. The GOTSim model includes a feature extraction module and a similarity calculation module. The feature extraction module is used to extract the feature representation of graph nodes, and the similarity calculation module is used to calculate the similarity between the customer's hair styling behavior activity graph and the standard hair styling behavior activity graph. The hair styling service content output module is used to output the hair styling service content corresponding to the standard hair styling behavior activity graph with the highest similarity to the customer's hair styling behavior activity graph.
[0038] Further, the feature extraction module in step S42 is composed of a multi-layer GCN network and is used to extract the features of the nodes of the hair styling behavior activity graph. The feature extraction steps are as follows:
[0039] S421: The service behavior names are saved in the nodes of the hair styling behavior activity graph. Perform one-hot encoding on all service behavior names to obtain the one-hot vector representation of the graph nodes.
[0040] S422: Respectively put the one-hot vectors of the nodes of the customer's hair styling behavior activity graph G1 and the standard hair styling behavior activity graph G2 (for better describing the calculation process of the feature extraction module and the similarity calculation module of GOTSim, the actual hair styling behavior activity graph and the standard hair styling behavior activity graph in steps S422 to S426 are tentatively set as G1 and G2) and the graph adjacency matrix constructed based on the edge set of the graph into the K-layer GCN network to obtain the graph node embedding matrix H G1,k and H G2,k , where H G,k ={h G,k,i |v i ∈V G , i = 1,…, N}, h G,k,i is the embedding vector of the node v i output by the k-th layer of GCN, V G is the node set of graph G, N is the number of elements in V G , and the calculation formula of h G,k,i is: where σ(·) is a non-linear activation function, d i is the degree of node v i plus 1, N(i) is the first-order neighborhood of node v i , and {W k , b k} represents the convolution kernel parameters of the k-th layer of GCN, which are shared by all nodes;
[0041] The similarity calculation module described in step S42 is used to calculate the similarity between two input graphs. The specific calculation steps are as follows:
[0042] S423: Calculate the similarity matrix of the k-th layer graphs G1 and G2 As follows:
[0043]
[0044] In the above matrix, is the cosine distance between node v in graph G1 i and node v in graph G2 j ; N1 is the number of nodes in graph G1, and N2 is the number of nodes in graph G2; Let h k,d and h k,a be the global embedding vectors converged during the training process of the GOTSim model. Then the deletion cost of node v i in G1 The insertion cost of node v j in G2
[0045] S424: Solve the optimal conversion cost from the k-th layer graph G1 to G2 where M is 's permutation matrix; After calculating the optimal conversion cost of each layer of GCN, the finally output similarity score is calculated as follows:
[0046] The training process of the GOTSim model is as follows:
[0047] S425: Use the haircut behavior activity graph training dataset as the input for model training, and set the number of training epochs (k e ) and the learning rate η; The purpose of training is to update the model's parameters {θ};
[0048] S426: The model training process is described as follows:
[0049] For each training epoch: Take m graph pairs for minibatch sampling: minibatch = {(G1, G2) l |l = 1, …, m}; For each GCN layer k in each graph pair (G1, G2), calculate the graph node embeddings of graphs G1 and G2 through the GOTSim feature extraction module and Calculate the similarity matrix of graphs G1 and G2 through the GOTSim similarity calculation module and the optimal conversion cost from the k-th layer graph G1 to G2 and 's permutation matrix Calculation For the gradient of the model parameters: Calculate using each layer Calculate
[0050] Calculate the minibatch loss:
[0051] Update the model parameters:
[0052] For the identity ID of id cus_now The steps for identifying the haircut service content are as follows:
[0053] S427: The customer haircut behavior activity diagram with the identity ID of id cus_now is The haircut service content output module first calculates the graph similarity with each standard haircut behavior activity diagram in the standard haircut behavior activity diagram library through the GOTSim model, and uses to save the maximum similarity calculated currently, and uses index G_max to save the index of the standard haircut behavior activity diagram participating in the calculation when the maximum similarity is obtained; after the iteration is completed, the haircut service content corresponding to the index G_max th standard haircut behavior activity diagram is output.
[0054] The beneficial effects of the present invention are as follows:
[0055] The present invention proposes a method for identifying haircut service behaviors, which uses a face recognition model and an object detection model to detect the customer's face and haircut tools in a surveillance video, and quickly and accurately identifies the customer receiving the service and the haircut behavior in the video frame according to the customer appearing in the specified area and the type of the haircut tool with the shortest Euclidean distance from the customer; at the same time, the present invention proposes a method for identifying haircut service content based on a behavior activity diagram, which constructs a haircut behavior activity diagram, effectively extracts the features of the behavior activity diagram, thereby calculating the similarity between the standard haircut behavior activity template diagram and the haircut behavior activity diagram in the actual scenario, and finally completes the identification of the customer's haircut service content.
[0056] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the accompanying drawings, where:
[0058] Figure 1 It is a flowchart of a method for identifying haircut service content based on a behavior activity diagram disclosed in an embodiment of the present invention;
[0059] Figure 2 It is a diagram of the installation positions of barbershop cameras, where (a) is a front view, (b) is a top view, and (c) is a side view;
[0060] Figure 3 It is a process for establishing a customer haircut behavior activity diagram;
[0061] Figure 4 It is a schematic diagram of a haircut behavior activity diagram. Specific Embodiments
[0062] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0063] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be construed as limitations on the present invention; for better illustrating the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0064] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as limitations on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0065] Such as Figure 1As shown in the figure, the present invention provides a method for identifying the content of hair-cutting services based on a behavior activity diagram, including the following steps:
[0066] S1: Set the relative positions of the haircut seats and the video monitoring devices, and obtain the video data stream of the video monitoring devices in real time; establish a barber face image library, and construct the corresponding relationship between the barber face images and the barber identities for identifying the barber faces in the video frames; establish a haircut tool image library, and construct the corresponding relationship between the haircut tools and the haircut behaviors for detecting and identifying the haircut tools in the video frames; establish a standard haircut behavior activity diagram library, and construct the corresponding relationship between the standard haircut behavior activity diagrams and the content of the haircut services for identifying the content of the haircut services of the customers;
[0067] S11: Set the relative positions of the haircut seats and the monitoring cameras to meet the requirements of the overall beauty of the barbershop. At the same time, it is also necessary to ensure that when a customer is sitting on the haircut seat, no one can pass between the haircut seat and the monitoring camera, meeting the requirements of the detection and recognition accuracy of the haircut tools and faces; the specific installation position and the viewing angle range of the camera are as Figure 2 shown in (a)-(c) in the figure. The camera is located at the top of the haircut mirror. Let the vertical distance between the camera and the floor be Cam_ver (unit: meter), the horizontal distance from the haircut mirror be Cam_hor (unit: meter), the horizontal viewing angle range of the camera be Cam_h_angle (unit: degree), and the pitch viewing angle range of the camera be Cam_p_angle (unit: degree).
[0068] Optionally, Cam_ver is set to 2 meters, Cam_hor is set to 0.5 meters, Cam_h_angle is set to 90°, and Cam_p_angle is set to 70°.
[0069] Specify the hardware conditions of the monitoring camera to meet the requirements of the model accuracy. The present invention requires that the resolution of the camera is not less than 1080P, which can meet the requirements of face recognition and haircut tool detection; at the same time, in order to reduce the number of nodes and edges of the subsequent haircut behavior activity diagram without losing too much information, the present invention requires that the frame rate of the camera is 20FPS;
[0070] S12: Establish a barber face image library, where each barber face image is named after the corresponding barber identity name for identifying the barber faces in the video frames; establish a haircut tool image library, where each haircut tool image is named after the corresponding service behavior name for identifying the service behaviors in the video frames, and the service behaviors and the haircut tools are in a one-to-many relationship;
[0071] S13: Establish a library of standard haircut activity diagrams (hereinafter collectively referred to as standard diagrams): A haircut service content consists of several service behaviors. The standard haircut activity diagram constructed by these service behaviors is represented as: G barc_i = <V barc_i , E barc_i >, where G barc_i is an undirected graph corresponding to the barc_i-th haircut service content, barc_i ∈ [0, barc_n - 1], and barc_n is the number of types of haircut service content; V barc_i is the set of service behavior nodes of G barc_i : V barc_i = {v barc_i_x | barc_i_x ∈ [0, barc_i_n - 1]}, barc_i_n is the number of elements in V barc_i , and each element in V barc_i stores the name of a service behavior; The edge set of G barc_i : E barc_i = {(v barc_i_x , v barc_i_y ) | barc_i_x, barc_i_y ∈ [0, barc_i_n - 1]}, and each element in E barc_i indicates that the service behavior corresponding to v barc_i_x occurs before the service behavior corresponding to v barc_i_y ; In actual storage, the storage fields of the standard diagram are divided into "label" and "graph". The "label" field stores the service behavior names of each node in the graph, corresponding to the vertex set V barc_i ; The "graph" field stores the index values of the endpoints of each edge in the graph in "label", corresponding to the edge set E barc_i ; The steps for establishing the standard diagram are as follows:
[0072] S131: According to the standard service process description corresponding to each haircut service content in the barbershop, shoot the service demonstration video of each haircut service content under the same conditions and perform frame extraction on the video;
[0073] S132: Mark the service behavior of each frame in the demonstration video corresponding to the barc_i-th haircut service content according to time. According to the types of marked service behaviors, add the service behavior nodes to the set V barc_i of service behavior nodes of G barc_i . At the same time, for the service behavior node v barc_i_x corresponding to a certain frame and the service behavior node v barc_i_y corresponding to the next frame, add (v barc_i_x , v barc_i_y ) to the edge set E barc_i of G barc_iFinally, the complete activity diagram G of the service demonstration video is formed. barc_i ;
[0074] It should be noted that during the process of annotating each frame of service behavior, for each service behavior annotation, it is necessary to first determine whether there is a node in V barc_i that stores the name of the service behavior. If not, a new node is added to V barc_i , and at the same time, the new node is made to store the name of the currently annotated service behavior. If it exists, no new node is added. Furthermore, if the newly added node is the first node in V barc_i , no edge is added to E barc_i ;
[0075] S133: Use the method described in step S132 to establish corresponding standard diagrams for all hair-cutting service contents to form a standard diagram library.
[0076] S2: Set a fixed detection range for video frames, and use existing face recognition and target tracking models to confirm the identities of the barber and the customer receiving the hair-cutting service; use a target detection model to detect and identify the hair-cutting tools that appear in the video frames, and mark their service behaviors according to the corresponding relationship with the service behaviors;
[0077] S21: Set a certain range where the face detection frame appears in the video frame after the customer sits down as the fixed detection frame range for service content video detection. The fixed detection frame is represented as: Frame det =(x det , y det , w det , h det ), where x det and y det respectively represent the horizontal and vertical axis coordinates of the upper left corner point of the fixed detection frame, and w det and h det respectively represent the number of pixels of the fixed detection frame along the vertical and horizontal axes, that is, the width and length of the fixed detection frame;
[0078] S22: Optionally, ArcFace face recognition model and MOTDT target tracking model are used to track faces in surveillance videos, wherein ArcFace is used to detect faces in video frames and identify the barber’s face, and MOTDT is used to associate the face frames output by ArcFace in each video frame to track faces. The specific steps for face tracking are as follows: the ArcFace face recognition model is pre-trained offline using the barber’s face image library, so that the ArcFace model can extract the position and features of the face in each frame, and the successfully identified barber’s face is assigned to the target video frame. The barber's identity name is obtained, and unrecognized faces are assigned non-duplicate customer IDs; the features of all faces extracted by ArcFace in the current frame are saved; the Kalman filter in the MOTDT target tracking model is used to predict the position of the face in the current video frame, and the prediction result and the face recognition result of the ArcFace model are updated through the MOTDT similarity calculation module and data association module to update the position of the face frame in the current frame, and the face features of the current frame are calculated and matched with the face features saved in the last face tracking, and the same identity information is assigned to each matched face to achieve face tracking;
[0079] The process of using the barber face image library to perform offline pre-training on the ArcFace face recognition model in step S22 is as follows:
[0080] S221: Use ResNet50 to extract facial features of each face image in the barber face image library face , L2 normalized face feature fea face and the weight W of the final fully connected layer arc And do the dot product of the two to get the cosine distance between the two, and then use the inverse cosine function to get the fea face With W arc The angle θ between arcface , an additive angle margin m arc Add to θ arcface On the top, calculate cos(θ arcface +m arc ) Obtain the predicted value of the target and scale it, and finally calculate the Softmax loss for the scaled predicted value and update the model parameters; continue to train and tune, and finally correctly output the face identity probability;
[0081] The process of using the MOTDT target tracking model to track the face in step S22 is as follows:
[0082] S222: First, the object classifier MTCNN in the ArcFace model outputs the face detection box of each frame and the probability that the face is confirmed. The Kalman filter in the MOTDT model estimates the position of each face prediction box in the current frame, and then calculates the trajectory confidence to evaluate the accuracy of the Kalman filter using temporal information; merge the face detection box and the face prediction box, calculate the unified score of the face merge box according to the face probability and the trajectory confidence, and input the face merge box and the corresponding unified score into the NMS algorithm to obtain the face candidate box of the current frame, which is used to remove unreliable and redundant boxes in the face merge box; input the filtered face candidate box into the ResNet50 network in the ArcFace model to identify the face identity information; perform IOU data association between the face box tracked last time and the face candidate box in the current frame, and at the same time calculate and match the face feature in the current frame with the face feature saved in the previous face tracking to achieve the tracking of the face;
[0083] S23: Considering that there may be multiple customer faces in the video frame, it is necessary to calculate the customer detection box within the fixed detection box range generated in each frame in step S22. This customer is confirmed as the customer receiving the haircut service. The specific calculation steps are as follows: Let the video frame in which the face is first detected be the first frame, then the ci-th customer detection box in the t-th frame is expressed as: Frame cus_t_ci =(x cus_t_ci , y cus_t_ci , w cus_t_ci , h cus_t_ci , id cus_t_ci ), where ci ∈ [0, cus_t_cn - 1], cus_t_cn is the total number of customer detection boxes in the t-th frame obtained in step S22, x cus_t_ci and y cus_t_ci represent the horizontal and vertical coordinates of the upper left corner point of Frame cus_t_ci ; w cus_t_ci and h cus_t_ci represent the width and length of Frame cus_t_ci ; id cus_t_ci represents the customer ID of Frame cus_t_ci ; it can be obtained that the center coordinates of Frame cus_t_ci are If and then confirm that the current customer is the customer receiving the haircut service, and save the customer ID of the customer detection box confirmed in the t-th frame and the customer detection box index of this box: id cus_t = id cus_t_ci , index cus_t = cus_t_ci(id cus_t and index cus_tThe initial value is empty), which is used for subsequent calculations of service behaviors associated with this customer;
[0084] S24: Optionally, use the barber tool image library to perform offline pre-training on the YoloV3 object detection model, which is used to output the position of the barber tool in each frame and label it with the corresponding service behavior name; as described in step S23, let the video frame in which the face is first detected be the first frame, then the bi-th barber tool detection box in the t-th frame is represented as Frame bart_t_bi =(x bart_t_bi ,y bart_t_bi ,w bart_t_bi ,h bart_t_bi ,p bart_t_bi ), where bi ∈ [0, bart_t_bn-1], and bart_t_bn is the total number of barber tool detection boxes in the t-th frame; x bart_t_bi and y bart_t_bi represent the horizontal and vertical axis coordinates of the upper left corner point of Frame bart_t_bi ; w bart_t_bi and h bart_t_bi represent the width and length of Frame bart_t_bi ; p bart_t_bi represents the service behavior name corresponding to Frame bart_t_bi .
[0085] The process of using the barber tool image library to perform offline pre-training on the YoloV3 object detection model in step S24 is as follows:
[0086] S241: First, the YoloV3 model scales the input size of the barber tool image, inputs the scaled image into the Darknet-53 network for feature extraction, which is used for predicting the border and category of the barber tool in the image; finally, calculate the object confidence loss, object category loss, and object localization loss for the model prediction results, which are used for updating the model network parameters; continuously train and optimize to enable the YoloV3 model to correctly output the position and category of the barber tool in each frame of the image;
[0087] Optionally, the input size of the barber tool image can be scaled to 416*416;
[0088] S31: Similar to step S23, let the video frame in which the face is first detected be the first frame; let the identity ID of the first confirmed customer after the first frame be id cus_now , starting from the frame where the identity ID of the customer is first confirmed as id cus_now , compare the saved id cus_t in each frame with id cus_now until the frame where id cus_t is empty for a certain number of consecutive frames or the non-empty id cus_t is not equal to idcus_now When the frame time ends, calculate the Euclidean distance between the center of each barber tool detection box in each frame and the center of the customer detection box with the identity ID of id cus_now Finally, obtain the service behavior name of the barber tool detection box participating in the calculation when the Euclidean distance is the smallest, and identify it as the service behavior associated with each frame of the customer with the identity ID of id cus_now ;
[0089] The steps for associating service behaviors are as follows:
[0090] S311: The center coordinates of the bi-th barber tool detection box Frame in the t-th frame bart_t_bi are The customer detection box with the confirmed identity ID of id in the t-th frame is represented as: cus_now where index
[0091]
[0092] represents the index of the customer detection box with the identity ID of id in the customer detection boxes of the t-th frame; cus_now_t and cus_now represent the horizontal and vertical axis coordinates of the upper left corner point; and represent the width and length; and represent ; The center coordinates of are further obtained as bart_t_bi the Euclidean distance between the center of Frame in the t-th frame and the center of ;
[0093]
[0094] S312: Let the minimum Euclidean distance calculated in the t-th frame be d t_min , and the index of the barber tool detection box participating in the calculation corresponding to it be index bart_t (d t_min is initially infinite, and index bart_t is initially empty); According to the calculation method of S311, loop to calculate the Euclidean distance between the center of each barber tool detection box in the t-th frame and until all barber tool detection boxes in the t-th frame are traversed; During the loop, every time an Euclidean distance is obtained, compare it with the current value of d t_min . If it is less than d t_min , then update d t_min to the currently obtained Euclidean distance, and at the same time update index bart_tIt is also updated to the barber tool detection frame index of the barber tool detection frame participating in the current calculation at the t-th frame; when the loop calculation ends, save the barbering behavior label of the index-th barber tool detection frame in the t-th frame bart_t barbering behavior label of the bart_t -th barber tool detection frame: Complete the recognition of the barbering behavior associated with the customer with the identity ID of id cus_now in the t-th frame;
[0095] Optionally, the initial value of d t_min can be set to 999999;
[0096] S32: According to the service behaviors recognized in each frame with the confirmed identity ID of id cus_now establish a barbering behavior activity graph of the customer with the identity ID of id cus_now hereinafter referred to as the customer graph), as the input of the barbering service content recognition model; the barbering behavior activity graph of the customer with the identity ID of id cus_now is represented as: where is an undirected graph, is the set of service behavior nodes of id cus_now _n is the number of elements in and each element in stores a service behavior name; is the edge set of Each element in indicates that the service behavior corresponding to occurs before the service behavior corresponding to In actual storage, the storage fields of the customer graph are divided into "label" and "graph". The "label" field stores the barbering behavior label of each node in the graph, corresponding to the vertex set
[0097] S321: Starting from the frame where the customer identity ID is first confirmed to be id cus_now compare the id cus_t saved in each frame with id cus_now Skip the frames where id cus_t is empty until id cus_t is not equal to id cus_now or the number of skipped frames exceeds a certain number and then end. Traverse the service behavior name p t_min recognized in each frame, and add the service behavior node to Service behavior node set Meanwhile, for the service behavior node corresponding to a certain frame and the behavior node corresponding to the next frame Add to Edge set Finally, a customer graph with customer identity ID as id cus_now is formed
[0098] Among them, let the number of consecutive frames with id cus_t being empty be t_null_cnt, and its upper limit be t_null_top. Optionally, let t_null_top be 18000, that is, the number of frames in a 15-minute video at a frame rate of 20 FPS;
[0099] It should be noted that during the process of identifying the service behavior of each frame, for each identified service behavior, it is necessary to first determine whether there is a node in it that stores the name of the service behavior. If not, a new node is added to and at the same time, let this node store the name of the currently labeled service behavior; if it exists, no new node is added. Furthermore, if the currently added node is the first node in, no edge is added to ;
[0100] S41: Construct a training dataset for the haircut behavior activity graph for offline pre-training of the service content recognition model. The specific construction process of the dataset is as follows: Obtain multiple customer graphs in the actual haircut scenario through steps S2 - S3, combine each obtained customer graph with each standard graph in the standard graph library to form multiple graph pairs, calculate the graph edit distance (GED) between the customer graph and the standard graph in the graph pair. GED is the minimum number of operations required to complete the mutual conversion between two graphs through the insertion, deletion, and replacement operations of points and edges. Considering accuracy, the present invention uses the A* algorithm to calculate GED; Store the "label", "graph" fields of each graph in each graph pair and their GED values as a json file. One json file serves as one piece of training data, and all training data constitutes the training dataset for the haircut behavior activity graph;
[0101] Optionally, according to the number of types n of the standard graphs, the number of training data can be set to 100 * n * n;
[0102] S42: Construct a hair salon service content recognition model for identifying the hair salon service content enjoyed by customers; the hair salon service content recognition model includes a GOTSim graph similarity calculation model and a hair salon service content output module. The GOTSim model is further divided into a feature extraction module and a similarity calculation module. The feature extraction module is used to extract the feature representation of graph nodes, and the similarity calculation module is used to calculate the similarity between the customer graph and the standard graph; the hair salon service content output module is used to output the hair salon service content corresponding to the standard graph with the highest similarity to the customer graph.
[0103] The feature extraction module described in step S42 is composed of a multi-layer GCN network and is used to extract the features of the hair salon behavior activity graph nodes. The feature extraction steps are as follows:
[0104] S421: The service behavior names are saved in the nodes of the hair salon behavior activity graph. One-hot encoding is performed on all service behavior names to obtain the one-hot vector representation of the graph nodes;
[0105] S422: (To better describe the calculation process of the feature extraction module and the similarity calculation module of GOTSim, the customer graph and the standard graph in steps S422 to S426 are temporarily designated as G1 and G2 respectively) The node one-hot vectors of the customer graph G1 and the standard graph G2 and the graph adjacency matrix constructed according to the edge set of the graph are respectively put into the multi-layer GCN network;
[0106] Optionally, to balance training efficiency and accuracy, the number of GCN network layers is set to 3, the number of neurons is set to [128, 64, 32], the convolution kernel size is set to [25, 10, 4, 2], the convolution output channel number is set to [16, 32, 64, 128], the kernel size in the pooling operation is set to [3, 3, 3, 2], and the pooling strategy is set to "Max-Pooling".
[0107] Therefore, the graph node embedding matrix output by the k-th (k ∈ [1, 3]) layer can be obtained in sequence and where H G,k ={h G,k,i |v i ∈V G , i = 1, …, N}, h G,k,i is the embedding vector of the node v i output by the k-th layer of GCN, V G is the node set of graph G, N is the number of elements in V G , and the calculation formula of h G,k,i is: where σ(·) is a non-linear activation function, d i is the node vi Increment the degree by 1. N(i) is the first-order neighborhood of node v i . {W k , b k} represents the convolutional kernel parameters of the k-th layer of GCN, which are shared by all nodes;
[0108] The similarity calculation module described in step S42 is used to calculate the similarity between the two input graphs. The specific calculation steps are as follows:
[0109] S423: Calculate the similarity matrix of the k-th layer graphs G1 and G2 as follows:
[0110]
[0111] In the above matrix, is the cosine distance between node v in graph G1 i and node v in graph G2 j . N1 is the number of nodes in graph G1, and N2 is the number of nodes in graph G2. Let h k,d and h k,a be the global embedding vectors converged during the training process of the GOTSim model. Then the deletion cost of node v i in G1 The insertion cost of node v j in G2
[0112] S424: Solve the optimal transformation cost from the k-th layer graph G1 to G2 where M is the permutation matrix; After calculating the optimal transformation cost for each layer of GCN, the final output similarity score is calculated as follows:
[0113] The training process of the GOTSim model is as follows:
[0114] S425: Use the haircut behavior activity graph training dataset as the input for model training, and set the training epoch (k e ) and the learning rate η; The purpose of training is to update each parameter {θ} of the model;
[0115] Optionally, the training epoch is set to 10, and the learning rate η is set to 0.001;
[0116] S426: The model training process is described as follows:
[0117] For each training epoch:
[0118] Sample m graph pairs for minibatch: minibatch = {(G1, G2) l | l = 1, …, m};
[0119] For each graph pair (G1, G2):
[0120] For each GCN layer k:
[0121] Calculate the graph node embeddings of graphs G1 and G2 through the GOTSim feature extraction module and
[0122] Calculate the similarity matrix of graphs G1 and G2 and the optimal transformation cost from graph G1 to graph G2 at the k-th layer and the permutation matrix therein as well as
[0123] Calculate the gradients of the model parameters:
[0124] End the loop for the GCN layer;
[0125] Using what is calculated in each layer Calculate
[0126] End the loop for the minibatch graph pairs;
[0127] Calculate the minibatch loss:
[0128] Update the model parameters:
[0129] End the current training epoch;
[0130] Optionally, the number of graph pairs in the minibatch is set to 128;
[0131] For the hair styling service content recognition steps with identity ID being id cus_now are as follows:
[0132] S427: For the customer graph with identity ID being id cus_now is The hair styling service content output module first calculates the graph similarity with each standard graph in the standard graph library through the GOTSim model and saves the currently calculated maximum similarity during the iteration process using and uses index G_maxSave the index of the standard image participating in the calculation when the maximum similarity is reached; after the iteration is completed, output the hair styling service content corresponding to the G_max indexth standard image.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for identifying the content of a haircut service based on an activity graph, characterized in that: It includes the following steps: S1: Set the relative positions of the barber seats and the video surveillance devices; establish a barber face image library, a barber tool image library, and a standard barber behavior activity library for detecting and identifying the barber's identity, barber tools, and barber service content; S2: Confirm the identities of the barber and the customer through a face recognition model and a target tracking model; detect and identify the barber tools that appear in the video frames through a target detection model; mark their service behaviors according to the corresponding relationships with the service behaviors; S3: In each frame, calculate the Euclidean distance between the customer detection box and the barber tool detection box; mark the service behavior of each frame with the barber tool detection box when the distance is minimized according to the corresponding relationship between the barber tool and the service behavior; construct a customer barber behavior activity graph; S4: Construct and train a barber service content recognition model based on graph similarity, and match the customer barber behavior activity graph with the standard barber behavior activity graph to determine the service content received by the customer in the barbershop.
2. The method for identifying the content of hair-cutting services based on the activity graph according to claim 1, wherein: In the step S1, it includes the following specific steps: S11: Set the relative positions of the barber seats and the monitoring cameras, and specify the hardware conditions of the monitoring cameras to meet the requirements of model accuracy; S12: Establish a barber face image library, where each barber face image is named after the corresponding barber's identity name for identifying the barber face in the video frames; establish a barber tool image library, where each barber tool image is named after the corresponding service behavior name for identifying the service behavior in the video frames, and the service behavior and the barber tool have a one-to-many relationship; S13: Establish a standard haircut behavior activity library: A haircut service content consists of several service behaviors. The standard haircut behavior activity diagram constructed by these service behaviors is represented as: G barc_i = <V barc_i , E barc_i >>, where G barc_i is an undirected graph corresponding to the barc_i-th haircut service content, barc_i ∈ [0, barc_n - 1], and barc_n is the number of types of haircut service content; V barc_i is the service behavior node set of G barc_i : V barc_i = {v barc_i_x | barc_i_x ∈ [0, barc_i_n - 1]}, barc_i_n is the number of elements in V barc_i , and each element in V barc_i stores the name of a service behavior; E barc_i is the edge set of G barc_i : E barc_i ={(v barc_i_x , v barc_i_y ) | barc_i_x, barc_i_y ∈ [0, barc_i_n - 1]}, where each element in E barc_i represents that the service behavior corresponding to v barc_i_x occurs prior to the service behavior corresponding to v barc_i_y .
3. The method for identifying hair styling service content based on an activity graph according to claim 2, wherein: The steps for establishing the standard barber behavior activity library in step S13 are as follows: S131: According to the standard service process descriptions corresponding to each barber service content in the barbershop, shoot service demonstration videos for each barber service content under the same conditions; S132: Mark the service actions of each frame in the demonstration video corresponding to the barc_i-th haircut service content according to time. According to the types of marked service actions, add the service action nodes to the set V barc_i of service action nodes of G barc_i . At the same time, for the service action node v barc_i_x corresponding to a certain frame and the service action node v barc_i_y corresponding to the next frame, add (v barc_i_x , v barc_i_y ) to the edge set E barc_i of G barc_i ; Finally, form the complete standard haircut behavior activity diagram G barc_i ; S133: Use the method described in step S132 to establish corresponding standard barber behavior activity graphs for all barber service contents to form a standard barber behavior activity library.
4. The method for identifying haircut service content based on an activity graph according to claim 1, wherein: In the step S2, it includes the following steps: S21: Set a certain range where the face detection frame appears in the video frame after the customer is seated as the fixed detection frame range for service content video detection; the fixed detection frame is represented as: Frame det =(x det , y det , w det , h det ), where x det and y det respectively represent the horizontal and vertical axis coordinates of the upper left corner point of the fixed detection frame, and w det and h det respectively represent the number of pixels of the fixed detection frame along the vertical axis and the horizontal axis, that is, the width and length of the fixed detection frame; S22: Use the face recognition model and the target tracking model to track the faces in the monitoring video. The face recognition model is used to detect the faces in the video frames and identify the barber faces among them, and the target tracking model is used to associate the face boxes output by the face recognition model in each video frame to achieve face tracking. The specific implementation steps of face tracking are as follows: Use the barber face image library to perform offline pre-training on the face recognition model so that the face recognition model can extract the positions and features of the faces in each frame, and at the same time assign the recognized barber faces to the corresponding barber identity names, assign non-repeated customer IDs to the unrecognized faces, and save the features of all current faces; Use the target tracking model to first predict the position of the face in the current frame, and update the position of the face bounding box in the current frame through the similarity calculation module and data association module of the target tracking model with the face detection result of the face recognition model. Calculate and match the face features in the current frame with the face features saved in the previous face tracking, and assign the same identity information to each matched face to achieve face tracking; S23: Confirm the customer detection boxes within the fixed detection box range generated in step S22 for each frame. The specific calculation steps are as follows: Let the video frame in which the face is first detected be the first frame. Then, the $i$-th customer detection box in the $t$-th frame is represented as: Frame cus_t_ci =(x cus_t_ci ,y cus_t_ci ,w cus_t_ci ,h cus_t_ci ,id cus_t_ci ), where $i\in[0,cus\_t\_cn - 1]$, $cus\_t\_cn$ is the total number of customer detection boxes in the $t$-th frame obtained in step S22. $x cus_t_ci $ and $y cus_t_ci $ represent the horizontal and vertical coordinates of the upper left corner point of Frame cus_t_ci . $w cus_t_ci $ and $h cus_t_ci $ represent the width and length of Frame cus_t_ci . $id cus_t_ci $ represents the customer ID of Frame cus_t_ci . Therefore, the center coordinates of Frame cus_t_ci are If and , then confirm that the current customer is the customer receiving the haircut service, and save the customer ID of the customer detection box confirmed in the $t$-th frame and the customer detection box index of this box: $id cus_t =id cus_t_ci , index cus_t =cus\_t\_ci; S24: Use the barber tool image library to perform offline pre-training on the object detection model, so that the object detection model can output the positions of barber tools in each video frame and label them with corresponding service behavior names; similar to step S23, assume that the video frame in which the face is first detected is the 1st frame, then the bi-th barber tool detection box in the t-th frame is represented as: Frame bart_t_bi =(x bart_t_bi ,y bart_t_bi ,w bart_t_bi ,h bart_t_bi ,p bart_t_bi ), where bi ∈ [0, bart_t_bn - 1], and bart_t_bn is the total number of barber tool detection boxes in the t-th frame; x bart_t_bi and y bart_t_bi represent the horizontal and vertical axis coordinates of the upper left corner point of Frame bart_t_bi , w bart_t_bi and h bart_t_bi represent the width and length of Frame bart_t_bi , and p bart_t_bi represents the service behavior name corresponding to Frame bart_t_bi .
5. The method for identifying haircut service content based on an activity graph according to claim 4, characterized in that: In step S3, the following steps are specifically included: S31: Similar to step S23, set the video frame when the face is first detected as the first frame; set the identity ID of the first confirmed customer after the first frame as id cus_now , starting from the frame when the identity ID of the customer is first confirmed as id cus_now , compare the id saved in each frame with id cus_t until the frames with empty id cus_now are consecutive for a certain number or the frames with non-empty id cus_t is not equal to id cus_t . Calculate the Euclidean distance between the center of each haircut tool detection box and the center of the customer detection box with the identity ID of id cus_now in each frame. Finally, obtain the service behavior name of the haircut tool detection box involved in the calculation when the Euclidean distance is the smallest, and identify it as the service behavior associated with each frame of the customer with the identity ID of id cus_now ; cus_now S32: Based on the confirmed identity ID being id cus_now establish an activity diagram of the customer's haircut behavior identified in each frame with the identity ID being id cus_now as the input to the haircut service content recognition model; the activity diagram of the customer's haircut behavior with the identity ID being id cus_now is represented as: where is an undirected graph, is the set of service behavior nodes of: id cus_now _n is the number of elements in, and each element in stores the name of a service behavior; is the set of edges of: each element in represents that the corresponding service behavior occurs prior to the corresponding service behavior.
6. The method for identifying hair styling service content based on an activity graph according to claim 5, wherein: The service behavior recognition steps in step S31 are as follows: S311: The Frame of the bi-th haircut tool detection box in the t-th frame bart_t_bi has the center coordinates of The customer detection box with the confirmed identity ID of idcus in the t-th frame _now is represented as: where index cus_now_t represents the customer detection frame index of the customer detection frame with identity ID as id cus_now in the t-th frame; and represent the horizontal and vertical axis coordinates of the upper left corner point; and represent the width and length of; The center coordinates of are Furthermore, the Euclidean distance between the center of Frame bart_t_bi and the center of in the t-th frame can be obtained S312: Let the minimum Euclidean distance calculated in the t-th frame be d t_min , and the index of the haircut tool detection box participating in the calculation corresponding to it be index bart_t ; According to the calculation method provided in step S311, loop to calculate the Euclidean distance between the center of each haircut tool detection box in the t-th frame and the center until all the haircut tool detection boxes in the t-th frame are traversed; During the loop, each time an Euclidean distance is calculated, compare it with the current d t_min value. If it is less than d t_min , then d t_min is updated to the currently calculated Euclidean distance, and at the same time index bart_t is also updated to the index of the haircut tool detection box participating in the calculation in the t-th frame; When the loop calculation ends, save the service behavior name of the index bart_t -th haircut tool detection box in the t-th frame: Complete the recognition of the service behavior associated with the customer with the identity ID of id cus_now in the t-th frame.
7. The method for identifying hair styling service content based on an activity graph according to claim 6, characterized in that: The steps for establishing the customer haircut behavior activity diagram in step S32 include: From the frame when the customer identity ID is first confirmed as id cus_now start, save the id in each frame cus_t and compare it with id cus_now until the frame where id cus_t is empty for a certain number of consecutive frames or the id cus_t is not equal to id cus_now ends. Traverse the service behavior name p t_min identified in each frame, and add the service behavior node to the service behavior node set At the same time, for the service behavior node corresponding to a certain frame and the behavior node corresponding to the next frame add to the edge set Finally, form the customer haircut behavior activity diagram with the customer identity ID as id cus_now 8. The method for identifying hair styling service content based on an activity graph according to claim 1, wherein: In step S4, the following steps are specifically included: S41: Construct a training dataset for the haircut behavior activity diagram for offline pre-training of the haircut service content recognition model; the specific construction process of the dataset is as follows: Obtain multiple customer haircut behavior activity diagrams in a real haircut scenario through steps S2 - S3, combine each obtained customer haircut behavior activity diagram with each standard haircut behavior activity diagram in the standard haircut behavior activity diagram library to form multiple graph pairs, calculate the graph edit distance of each customer haircut behavior activity diagram and standard haircut behavior activity diagram in each graph pair using the A* algorithm, and finally combine the node set and edge set of each graph in each graph pair and the GED of the two graphs to form training data. All the training data constitutes the training dataset for the haircut behavior activity diagram; S42: Construct a haircut service content recognition model for identifying the haircut service content enjoyed by customers in the barbershop; the haircut service content recognition model includes a GOTSim graph similarity calculation model and a haircut service content output module. The GOTSim model includes a feature extraction module and a similarity calculation module. The feature extraction module is used to extract the feature representation of the graph nodes, and the similarity calculation module is used to calculate the similarity between the customer haircut behavior activity diagram and the standard haircut behavior activity diagram; the haircut service content output module is used to output the haircut service content corresponding to the standard haircut behavior activity diagram with the highest similarity to the customer haircut behavior activity diagram.
9. The method for identifying hair styling service content based on an activity graph according to claim 8, wherein: The feature extraction module described in step S42 is composed of a multi-layer GCN network and is used to extract the features of the nodes of the haircut behavior activity diagram. The feature extraction steps are as follows: S421: The service behavior names are saved in the nodes of the haircut behavior activity diagram. One-hot encoding is performed on all the service behavior names to obtain the one-hot vector representation of the graph nodes; S422: Respectively put the node one-hot vectors of the customer haircut behavior activity diagram G1 and the standard haircut behavior activity diagram G2 and the graph adjacency matrix constructed based on the edge set of the graph into the K-layer GCN network to obtain the graph node embedding matrix output by the k-th layer and where k ∈ [1, K], H G,k ={h G,k,i |v i ∈V G , i = 1,..., N}, h G,k,i is the embedding vector of the node v i output by the k-th layer GCN, V G is the node set of the graph G, N is the number of elements in V G , and the calculation formula of h G,k,i is: where σ(·) is a non-linear activation function, d i is the degree of the node v i plus 1, N(i) is the first-order neighborhood of the node v i , and {W k , b k} represents the convolution kernel parameters of the k-th layer GCN, which are shared by all nodes; The similarity calculation module described in step S42 is used to calculate the similarity between the two input graphs. The specific calculation steps are: S423: Calculate the similarity matrix of the k-th layer graphs G1 and G2 as follows: In the above matrix, is the cosine distance between node v in graph G1 i and node v in graph G2 j ; N1 is the number of nodes in graph G1, and N2 is the number of nodes in graph G2; Let h k,d and h k,a be the global embedding vectors obtained during the convergence of the GOTSim model training process. Then, the deletion cost of node v i in G1 The insertion cost of node v j in G2 S424: Solve the optimal conversion cost from graph G1 to graph G2 at the k-th layer where M is the permutation matrix; after calculating the optimal conversion cost for each layer of GCN, the final output similarity score is calculated as follows: The training process of the GOTSim model is as follows: S425: Use the haircut activity diagram training dataset as the input for model training, and set the number of training epochs (k e ) and the learning rate η; the purpose of training is to update the model parameters {θ}; S426: The model training process is described as follows: For each training epoch: Sample m graph pairs for minibatch: minbatch = {(G1, G2) l | l = 1,..., m}; For each GCN layer k in each graph pair (G1, G2), calculate the graph node embeddings of graphs G1 and G2 through the GOTSim feature extraction module and Calculate the similarity matrix of graphs G1 and G2 through the GOTSim similarity calculation module and the optimal transformation cost from graph G1 to graph G2 at the k-th layer and the permutation matrix among them Calculate For the gradients of the model parameters: Calculate using the calculated for each layer Calculate Calculate the minibatch loss: Update model parameters: The identity ID is id cus_now The steps for identifying the content of the haircut service are as follows: S427: The identity ID is id cus_now The haircut behavior activity diagram of the customer is The haircut service content output module first calculates the graph similarity with each standard haircut behavior activity diagram in the standard haircut behavior activity diagram library through the GOTSim model, and uses to save the maximum similarity calculated currently, and uses index G_max to save the index of the standard haircut behavior activity diagram participating in the calculation when the maximum similarity is obtained; after the iteration is completed, the haircut service content corresponding to the index G_max th standard haircut behavior activity diagram is output.
Citation Information
Patent Citations
Intelligent self-selection type consumption managing tool for hot pot restaurants
CN103810790A
Self-service intelligent haircut system and method
CN114170732A