Digital human interaction method and device based on large language model
By using a digital human interaction method based on a large language model in police alarms, and using multimodal compression model and deep neural network for data processing, the problem of difficulty in arranging large language models and slow information transmission speed on the mobile phone is solved, and efficient data transmission and decoding effects are achieved.
Patent Information
- Application Number
- CN202411941771.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, in police alarms, it is difficult for mobile phones to arrange large language models, and the cloud information transmission speed is slow, especially the video information transmission time is poor.
The digital human interaction method based on the large language model is adopted, and data compression and decompression is carried out through the multimodal compression model of the client and the cloud, and feature extraction and decoding is used for deep neural networks to achieve efficient data transmission and decoding.
The balance between high compression rate and high restore quality is achieved, the storage volume and time of data transmission is reduced, the efficiency of information transmission is improved, and the contradiction between compression rate and restore quality in the prior art is better than the contradiction between compression rate and restore quality.
Smart Images

Figure CN119961429A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language models, and in particular to a digital human interaction method based on a large language model and a digital human interaction device based on a large language model. Background Art
[0002] When making a police call, the caller usually transmits multiple information, such as text messages, voice messages, and video information (such as the situation at the crime scene). The terminals usually used by customers are mobile phones. In the existing technology, the layout of large models requires a large space and high computing power, and it is unlikely that mobile phones can be equipped with large models. It is unrealistic to equip each mobile phone with a large model. However, if the cloud is used for communication, the information transmission speed of the existing technology is slow, especially when there is video and the video is large.
[0003] Therefore, it is desired to have a technical solution to overcome or at least alleviate at least one of the above-mentioned defects of the prior art.
[0004] Application Contents
[0005] The purpose of the present application is to provide a digital human interaction method based on a large language model to overcome or at least alleviate at least one of the above-mentioned defects of the prior art.
[0006] To achieve the above-mentioned purpose, the present application provides a digital human interaction method based on a large language model for police alarm, and the digital human interaction method based on a large language model includes:
[0007] The client obtains the voice information to be interacted, the video information to be interacted, and the text information to be interacted input by the user;
[0008] The client obtains the trained client multimodal compression model;
[0009] The client compresses the voice information to be interacted, the video information to be interacted, and the text information to be interacted by using the trained multimodal compression model to obtain first compressed information and sends the first compressed information to the cloud;
[0010] The cloud decompresses the obtained first compressed information through a trained cloud multimodal compression model, thereby obtaining original voice information to be interacted, original video information to be interacted, and original text information to be interacted;
[0011] Obtain a large, trained language model in the cloud;
[0012] The cloud inputs the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted into a trained large language model to obtain reply information and reply video information;
[0013] The cloud compresses the obtained reply information and reply video information through the cloud multimodal compression model to obtain second compressed information;
[0014] The cloud sends the second compressed information to the client.
[0015] Optionally, the client multimodal compression model and the cloud multimodal compression model are the same model and have the same hyperparameters;
[0016] The digital human interaction method based on the large language model further comprises:
[0017] Obtain a multimodal compression model;
[0018] Training a multimodal compression model;
[0019] The trained multimodal compression model is deployed to the cloud and the client, wherein the multimodal compression model located on the client is the client multimodal compression model, and the multimodal compression model located on the cloud is the cloud multimodal compression model.
[0020] Optionally, the multimodal compression model includes a spatiotemporal feature extractor, a verifier and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and spatiotemporal features from voice information, video information and text information, the verifier is used to receive visual spatial features and spatiotemporal features as input, obtain recognition results through spatial features and spatiotemporal features, and the decoder is used to receive spatiotemporal features as input, and decode the spatiotemporal features back to the original voice information, video information and text information;
[0021] The training of the multimodal compression model includes:
[0022] Acquire a training data set, wherein the training data set includes multiple sets of training data, each set of training data includes at least one voice information, video information, and text information;
[0023] Initialize the model parameters and training parameters of the multimodal compression model;
[0024] Inputting the training data set into the spatiotemporal feature extractor to obtain spatial features and spatiotemporal features;
[0025] Inputting the spatial features and the spatiotemporal features into a verifier to obtain a verifier classification result;
[0026] Obtaining a first loss function according to the verifier category result;
[0027] Input the spatiotemporal features into the decoder to obtain the original voice information, video information, and text information;
[0028] Obtaining a second loss function according to the original voice information, video information, and text information;
[0029] Obtain a multi-task joint loss function according to the first loss function and the second loss function;
[0030] Repeat the above steps until the iteration is completed or the loss value tends to be stable, then the training is completed, and the model parameters at the time of training completion are obtained.
[0031] Optionally, the cloud inputs the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted into a trained large language model to obtain reply information and reply video information, including:
[0032] Fusing the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted, thereby obtaining a fusion feature;
[0033] The fused features are input into the large language model to obtain reply information and reply video information.
[0034] Optionally, after the client obtains the voice information to be interacted, the video information to be interacted, and the text information to be interacted input by the user, the digital human interaction method based on the large language model further includes:
[0035] The client obtains the video semantic recognition model and the text semantic recognition model;
[0036] The client recognizes semantic information of the voice information to be interacted with according to the voice information to be interacted with and the text semantic recognition model;
[0037] The client identifies semantic information of the text information to be interacted with according to the text information to be interacted with and the text semantic recognition model;
[0038] The client obtains the semantic information of the video to be interacted with according to the video information to be interacted with and the video semantic recognition model;
[0039] The client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the semantic information of the video to be interacted.
[0040] Optionally, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the semantic information of the video to be interacted, including:
[0041] Obtain a relational knowledge graph, wherein the relational knowledge graph includes a plurality of nodes and a node distance between each node and other nodes, and each node represents a preset semantic information;
[0042] Respectively obtaining nodes corresponding to the semantic information of the voice information to be interacted, nodes corresponding to the semantic information of the text information to be interacted, and nodes corresponding to the semantic information of the video information to be interacted;
[0043] It is determined whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes.
[0044] Optionally, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, including:
[0045] When the number of nodes corresponding to the semantic information of the voice information to be interacted, the number of nodes corresponding to the semantic information of the text information to be interacted, and the number of nodes corresponding to the semantic information of the video to be interacted are all one, the sum of the node distances between the nodes is obtained. When the sum of the node distances between the nodes is less than a first preset threshold, it is determined that the obtained node distances between the nodes determine that the obtained voice information to be interacted, the video information to be interacted, and the text information to be interacted need to be transmitted to the client at the same time.
[0046] Optionally, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, further comprising:
[0047] When the number of nodes corresponding to the semantic information of the voice information to be interacted is one and the number of nodes corresponding to the semantic information of the text information to be interacted and the number of nodes corresponding to the semantic information of the video information to be interacted are more than one, the following method is used to determine whether to transmit them to the client at the same time:
[0048] Obtaining the node distances between the nodes corresponding to the semantic information of each speech information to be interacted and the nodes of the semantic information of each text information to be interacted, the node distances being referred to as first node distances;
[0049] Obtaining the node distances between the nodes corresponding to the semantic information of each to-be-interacted voice information and the nodes corresponding to the semantic information of each to-be-interacted video information, the node distances being referred to as second node distances;
[0050] When the sum of a first node distance and a second node distance is less than a first preset threshold, the node distances between the acquired nodes are determined to determine whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
[0051] Optionally, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, further comprising:
[0052] When the number of nodes corresponding to the semantic information of the voice information to be interacted, the number of nodes corresponding to the semantic information of the text information to be interacted, and the number of nodes corresponding to the semantic information of the video information to be interacted are all multiple, the following method is used to determine whether to transmit them to the client at the same time:
[0053] A scatter plot is made with any node as the origin. All nodes are located on the scatter plot, and the shortest straight line distance between each node and the origin is the node distance between the node and the origin.
[0054] Cluster each node on the scatter plot to obtain clusters;
[0055] When each node in a cluster has at least one semantic information belonging to the voice information to be interacted, at least one node corresponding to the semantic information belonging to the text information to be interacted, and at least one node belonging to the semantic information of the video to be interacted, then the node distance between the obtained nodes is determined to determine whether the obtained voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
[0056] The present application also provides a digital human interaction device based on a large language model, a client and a cloud of the digital human interaction device based on a large language model, and the client and the cloud cooperate to implement the digital human interaction device based on the large language model as described above.
[0057] The digital human interaction method based on a large language model of the present application uses the compression model of the present application during data transmission and decoding, which can compress the transmission information and compress the original data into a feature vector, greatly reducing the volume of stored data and achieving a high compression rate. At the same time, by combining the joint training of the decoder module and the verifier, the model parameters are optimized so that the generated feature vector can effectively retain the key visual details of the original data, achieving a balance between extremely high compression rate and high restoration quality, which is better than the contradiction between compression rate and restoration quality in the prior art that is difficult to reconcile. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1It is a flowchart of a digital human interaction method based on a large language model according to an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical scheme and advantages of the implementation of this application clearer, the technical scheme in the embodiment of this application will be described in more detail below in conjunction with the drawings in the embodiment of this application. In the drawings, the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The described embodiments are part of the embodiments of this application, not all of them. The embodiments described below with reference to the drawings are exemplary and are intended to be used to explain this application, and should not be construed as limitations on this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. The embodiments of this application are described in detail below in conjunction with the drawings.
[0060] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the scope of protection of the present application.
[0061] like Figure 1 The digital human interaction method based on the large language model is used for police alarm, and the digital human interaction method based on the large language model includes:
[0062] Step 1: The client obtains the voice information to be interacted, the video information to be interacted, and the text information to be interacted input by the user;
[0063] Step 2: The client obtains the trained client multimodal compression model;
[0064] Step 3: The client compresses the to-be-interacted voice information, the to-be-interacted video information, and the to-be-interacted text information through the trained multimodal compression model to obtain first compressed information and sends the first compressed information to the cloud;
[0065] Step 4: The cloud decompresses the first compressed information obtained by using a trained cloud multimodal compression model, thereby obtaining original voice information to be interacted, video information to be interacted, and text information to be interacted;
[0066] Step 5: Obtain the trained large language model in the cloud;
[0067] Step 6: The cloud inputs the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted into the trained large language model to obtain reply information and reply video information;
[0068] Step 7: The cloud compresses the obtained reply information and reply video information through the cloud multimodal compression model to obtain second compressed information;
[0069] Step 8: The cloud sends the second compressed information to the client.
[0070] The digital human interaction method based on a large language model of the present application uses the compression model of the present application during data transmission and decoding, which can compress the transmission information and compress the original data into a feature vector, greatly reducing the volume of stored data and achieving a high compression rate. At the same time, by combining the joint training of the decoder module and the verifier, the model parameters are optimized so that the generated feature vector can effectively retain the key visual details of the original data, achieving a balance between extremely high compression rate and high restoration quality, which is better than the contradiction between compression rate and restoration quality in the prior art that is difficult to reconcile.
[0071] In this embodiment, the client multimodal compression model and the cloud multimodal compression model are the same model and have the same hyperparameters;
[0072] The digital human interaction method based on the large language model further comprises:
[0073] Obtain a multimodal compression model;
[0074] Training a multimodal compression model;
[0075] The trained multimodal compression model is deployed to the cloud and the client, wherein the multimodal compression model located on the client is the client multimodal compression model, and the multimodal compression model located on the cloud is the cloud multimodal compression model.
[0076] The digital human interaction method based on the large language model of the present application has the following advantages:
[0077] 1. The digital human interaction method based on a large language model of the present application uses the compression model of the present application during data transmission and decoding, which can compress the transmission information and compress the original data into a feature vector, greatly reducing the volume of stored data and achieving a high compression rate. At the same time, by combining the joint training of the decoder module and the verifier, the model parameters are optimized so that the generated feature vector can effectively retain the key visual details of the original data, achieving a balance between extremely high compression rate and high restoration quality, which is better than the contradiction between compression rate and restoration quality in the prior art that is difficult to reconcile.
[0078] 2. This application uses a decoder module based on a deep neural network to decode feature vectors, which can use the good parallel computing capabilities of deep neural networks to further speed up the decoding speed of compressed feature vectors and improve the decoding efficiency compared to existing methods.
[0079] 3. This application proposes a strategy for multi-task joint training of the verifier and decoder. By optimizing the structure and training method of the deep generative model, the reconstruction effect of the detail area is significantly improved, thereby ensuring that the detail restoration effect under high compression rate is more realistic.
[0080] In this embodiment, the multimodal compression model includes a spatiotemporal feature extractor, a verifier, and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and spatiotemporal features from voice information, video information, and text information; the verifier is used to receive visual spatial features and spatiotemporal features as input, and obtain recognition results through spatial features and spatiotemporal features; the decoder is used to receive spatiotemporal features as input, and decode the spatiotemporal features back to the original voice information, video information, and text information;
[0081] The training of the multimodal compression model includes:
[0082] Acquire a training data set, wherein the training data set includes multiple sets of training data, each set of training data includes at least one voice information, video information, and text information;
[0083] Initialize the model parameters and training parameters of the multimodal compression model;
[0084] Inputting the training data set into the spatiotemporal feature extractor to obtain spatial features and spatiotemporal features;
[0085] Inputting the spatial features and the spatiotemporal features into a verifier to obtain a verifier classification result;
[0086] Obtaining a first loss function according to the verifier category result;
[0087] Input the spatiotemporal features into the decoder to obtain the original voice information, video information, and text information;
[0088] Obtaining a second loss function according to the original voice information, video information, and text information;
[0089] Obtain a multi-task joint loss function according to the first loss function and the second loss function;
[0090] Repeat the above steps until the iteration is completed or the loss value tends to be stable, then the training is completed, and the model parameters at the time of training completion are obtained.
[0091] In this embodiment, the spatiotemporal feature extractor includes a ConvNext network and an LSTM module. The verifier includes a Translayer, which receives spatial features and spatiotemporal features as input, and further fuses the spatial features and temporal features to obtain a recognition result through an attention mechanism similar to that of a Transformer. The decoder includes a Ground Decoding module, which receives spatiotemporal features as input, and decodes the spatiotemporal features back to video frame data in a manner similar to that of a Transformer Decoder.
[0092] In this embodiment, the verifier includes a normalization layer, a fully connected layer, a pixel normalization layer, and a convolutional layer.
[0093] In this embodiment, the first loss function is as follows:
[0094]
[0095] Among them, L class represents the loss value, y is the true label value of the training data, is the recognition result of the training data, n is the number of categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part are the verifier and spatiotemporal feature extractor.
[0096] In this embodiment, the second loss function is as follows:
[0097]
[0098] Among them, L recon represents the reconstruction loss value, y is the original training data, is the decompressed training data generated by the model, C is the number of training data categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part include the spatiotemporal feature extractor and decoder.
[0099] In this embodiment, the multi-task joint loss function is as follows:
[0100] L total =αL class +βL recon
[0101] Among them, α and β are weight parameters used to balance the two losses, and the entire model parameters, including the spatiotemporal feature extractor, verifier, and decoder, are optimized through back propagation.
[0102] In this embodiment, the cloud inputs the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted into a trained large language model to obtain reply information and reply video information, including:
[0103] Fusing the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted, thereby obtaining a fusion feature;
[0104] The fused features are input into the large language model to obtain reply information and reply video information.
[0105] In this embodiment, the fusing of the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted to obtain the fusion feature includes:
[0106] Preprocessing the to-be-interacted video information to obtain a first feature vector;
[0107] Performing a second preprocessing on the interactive voice information to obtain a second feature vector;
[0108] Preprocess the interactive text information to obtain the third eigenvector;
[0109] The first feature vector, the second feature vector and the third feature vector are fused through a regression mapping network to generate a fused feature.
[0110] Specifically, preprocessing the to-be-interacted video information to obtain a first feature vector includes:
[0111] Performing format adjustment on each frame of the interactive video information, thereby obtaining an image after the format adjustment;
[0112] Normalizing the images adjusted in various formats to form a normalized image;
[0113] Perform feature extraction on the normalized image to obtain a first feature vector.
[0114] The performing second preprocessing on the interactive voice information to obtain a second feature vector includes:
[0115] De-noise the voice information to be interacted with and perform voice recognition to obtain text information;
[0116] Perform word segmentation on text information to obtain token combinations;
[0117] The encoder is used to extract features from the token combination to obtain the second feature vector.
[0118] Preprocessing the interactive text information to obtain the third feature vector includes:
[0119] Perform word segmentation on text information to obtain token combinations;
[0120] The encoder is used to extract features from the token combination to obtain the third feature vector.
[0121] In this embodiment, after the client obtains the voice information to be interacted, the video information to be interacted, and the text information to be interacted input by the user, the digital human interaction method based on the large language model further includes:
[0122] The client obtains the video semantic recognition model and the text semantic recognition model;
[0123] The client recognizes semantic information of the voice information to be interacted with according to the voice information to be interacted with and the text semantic recognition model;
[0124] The client identifies semantic information of the text information to be interacted with according to the text information to be interacted with and the text semantic recognition model;
[0125] The client obtains the semantic information of the video to be interacted with according to the video information to be interacted with and the video semantic recognition model;
[0126] The client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the semantic information of the video to be interacted.
[0127] In this embodiment, the video semantic recognition model is a video behavior recognition model, such as an RBF neural network.
[0128] In this embodiment, both the text information to be interacted with and the voice information to be interacted with can be recognized using the text semantic recognition model. The difference is that the voice information needs to be converted into text information and then input into the text semantic recognition model.
[0129] In this embodiment, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the semantic information of the video to be interacted, including:
[0130] Obtain a relational knowledge graph, wherein the relational knowledge graph includes a plurality of nodes and a node distance between each node and other nodes, and each node represents a preset semantic information;
[0131] Respectively obtaining nodes corresponding to the semantic information of the voice information to be interacted, nodes corresponding to the semantic information of the text information to be interacted, and nodes corresponding to the semantic information of the video information to be interacted;
[0132] It is determined whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes.
[0133] In this embodiment, when a police alarm is actually carried out, some information (such as text information, video information and voice information) is usually transmitted to the cloud through the mobile phone every time period. The information in this time period may be relevant or irrelevant. For example, it may all be descriptions of a certain alarm situation, or some may be descriptions of the alarm situation, and some may be other things. This will result in a waste of cloud computing power if useless and irrelevant information is also transmitted to the cloud, and it may cause the large model on the cloud to be unable to correctly answer the questions that should be answered. Therefore, the present application uses the above method to determine whether the voice information to be interacted, the video information to be interacted, and the text information to be interacted obtained within a time period need to be transmitted to the client at the same time.
[0134] In this embodiment, the relational knowledge graph may include multiple nodes, and the distance between each node and other nodes actually represents the correlation. For example, the preset semantic information represented by one node is robbery, and the preset semantic information represented by another node is bleeding. The two may be associated with each other, so the node distance between the two nodes is relatively small. On the contrary, the preset semantic information represented by one node is robbery, and the preset semantic information represented by another node is sleeping, then the node distance between the two nodes is relatively large.
[0135] In this embodiment, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, including:
[0136] When the number of nodes corresponding to the semantic information of the voice information to be interacted, the number of nodes corresponding to the semantic information of the text information to be interacted, and the number of nodes corresponding to the semantic information of the video to be interacted are all one, the sum of the node distances between the nodes is obtained. When the sum of the node distances between the nodes is less than a first preset threshold, it is determined that the obtained node distances between the nodes determine that the obtained voice information to be interacted, the video information to be interacted, and the text information to be interacted need to be transmitted to the client at the same time.
[0137] When the sum of the node distances of the three is less than a first preset threshold, it is determined that the relationship between the three is relatively close. It can be understood that the first preset threshold can be set as needed.
[0138] In this embodiment, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, further comprising:
[0139] When the number of nodes corresponding to the semantic information of the voice information to be interacted is one and the number of nodes corresponding to the semantic information of the text information to be interacted and the number of nodes corresponding to the semantic information of the video information to be interacted are more than one, the following method is used to determine whether to transmit them to the client at the same time:
[0140] Obtaining the node distances between the nodes corresponding to the semantic information of each speech information to be interacted and the nodes of the semantic information of each text information to be interacted, the node distances being referred to as first node distances;
[0141] Obtaining the node distances between the nodes corresponding to the semantic information of each to-be-interacted voice information and the nodes corresponding to the semantic information of each to-be-interacted video information, the node distances being referred to as second node distances;
[0142] When the sum of a first node distance and a second node distance is less than a first preset threshold, the node distances between the acquired nodes are determined to determine whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
[0143] It is understandable that the above method can also be used when the semantic information of the text information to be interacted corresponds to one node and the number of the other two nodes exceeds one, and no further details will be given here.
[0144] It is understandable that the above method can also be used when the semantic information of the video semantic information to be interacted corresponds to one node and the number of the other two nodes exceeds one, and no further details are given here.
[0145] It can be understood that when two of the semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the semantic information of the video to be interacted are one node and the other one is multiple nodes, multiple triplets can be formed by respectively forming triplets (one node for the semantic information of the voice information to be interacted, one node for the semantic information of the text information to be interacted, and one node for the semantic information of the video to be interacted), and when the sum of the node distances of the three in each triplet is less than a first preset threshold, if so, the node distances between the obtained nodes are judged to determine whether the obtained voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
[0146] In this embodiment, the client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, further comprising:
[0147] When the number of nodes corresponding to the semantic information of the voice information to be interacted, the number of nodes corresponding to the semantic information of the text information to be interacted, and the number of nodes corresponding to the semantic information of the video information to be interacted are all multiple, the following method is used to determine whether to transmit them to the client at the same time:
[0148] A scatter plot is made with any node as the origin. All nodes are located on the scatter plot, and the shortest straight line distance between each node and the origin is the node distance between the node and the origin.
[0149] Cluster each node on the scatter plot to obtain clusters;
[0150] When each node in a cluster has at least one semantic information belonging to the voice information to be interacted, at least one node corresponding to the semantic information belonging to the text information to be interacted, and at least one node belonging to the semantic information of the video to be interacted, then the node distance between the obtained nodes is determined to determine whether the obtained voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
[0151] The present application also provides a digital human interaction device based on a large language model, a client and a cloud of the digital human interaction device based on a large language model, and the client and the cloud cooperate to implement the digital human interaction device based on the large language model as described above.
[0152] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application is described in detail with reference to the above embodiments, a person skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A digital human interaction method based on a large language model for police alarm, characterized in that: The digital human interaction method based on the large language model includes: The client obtains the voice information to be interacted, the video information to be interacted, and the text information to be interacted input by the user; The client obtains the trained client multimodal compression model; The client compresses the voice information to be interacted, the video information to be interacted, and the text information to be interacted by using the trained multimodal compression model to obtain first compressed information and sends the first compressed information to the cloud; The cloud decompresses the obtained first compressed information through a trained cloud multimodal compression model, thereby obtaining original voice information to be interacted, original video information to be interacted, and original text information to be interacted; Obtain a large, trained language model in the cloud; The cloud inputs the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted into a trained large language model to obtain reply information and reply video information; The cloud compresses the obtained reply information and reply video information through the cloud multimodal compression model to obtain second compressed information; The cloud sends the second compressed information to the client.
2. The digital human interaction method based on a large language model as claimed in claim 1, characterized in that: The client multimodal compression model and the cloud multimodal compression model are the same model and have the same hyper parameters; The digital human interaction method based on the large language model further comprises: Obtain a multimodal compression model; Training a multimodal compression model; The trained multimodal compression model is deployed to the cloud and the client, wherein the multimodal compression model located on the client is the client multimodal compression model, and the multimodal compression model located on the cloud is the cloud multimodal compression model.
3. The digital human interaction method based on a large language model as claimed in claim 2, characterized in that: The multimodal compression model includes a spatiotemporal feature extractor, a verifier and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and spatiotemporal features from voice information, video information and text information, the verifier is used to receive visual spatial features and spatiotemporal features as input, obtain recognition results through spatial features and spatiotemporal features, and the decoder is used to receive spatiotemporal features as input, and decode the spatiotemporal features back to the original voice information, video information and text information; The training of the multimodal compression model includes: Acquire a training data set, wherein the training data set includes multiple sets of training data, each set of training data includes at least one voice information, video information, and text information; Initialize the model parameters and training parameters of the multimodal compression model; Inputting the training data set into the spatiotemporal feature extractor to obtain spatial features and spatiotemporal features; Inputting the spatial features and the spatiotemporal features into a verifier to obtain a verifier classification result; Obtaining a first loss function according to the verifier category result; Input the spatiotemporal features into the decoder to obtain the original voice information, video information, and text information; Obtaining a second loss function according to the original voice information, video information, and text information; Obtain a multi-task joint loss function according to the first loss function and the second loss function; Repeat the above steps until the iteration is completed or the loss value tends to be stable, then the training is completed, and the model parameters at the time of training completion are obtained.
4. The digital human interaction method based on a large language model as claimed in claim 3, characterized in that: The cloud inputs the original voice information to be interacted, the original video information to be interacted, and the original text information to be interacted into the trained large language model to obtain reply information and reply video information, including: Fusing the original voice information to be interacted, the original video information to be interacted, and the text information to be interacted, thereby obtaining a fusion feature; The fused features are input into the large language model to obtain reply information and reply video information.
5. The digital human interaction method based on a large language model as claimed in claim 4, characterized in that: After the client obtains the voice information to be interacted, the video information to be interacted, and the text information to be interacted input by the user, the digital human interaction method based on the large language model further includes: The client obtains the video semantic recognition model and the text semantic recognition model; The client recognizes semantic information of the voice information to be interacted with according to the voice information to be interacted with and the text semantic recognition model; The client identifies semantic information of the text information to be interacted with according to the text information to be interacted with and the text semantic recognition model; The client obtains the semantic information of the video to be interacted with according to the video information to be interacted with and the video semantic recognition model; The client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the semantic information of the video to be interacted.
6. The digital human interaction method based on a large language model as claimed in claim 5, characterized in that: The client determines whether the acquired voice information to be interacted, the acquired video information to be interacted, and the acquired text information to be interacted need to be transmitted to the client at the same time according to the acquired semantic information of the voice information to be interacted, the semantic information of the text information to be interacted, and the acquired video semantic information to be interacted, including: Obtain a relational knowledge graph, wherein the relational knowledge graph includes a plurality of nodes and a node distance between each node and other nodes, and each node represents a preset semantic information; Respectively obtaining nodes corresponding to the semantic information of the voice information to be interacted, nodes corresponding to the semantic information of the text information to be interacted, and nodes corresponding to the semantic information of the video information to be interacted; It is determined whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes.
7. The digital human interaction method based on a large language model as claimed in claim 6, characterized in that: The client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, including: When the number of nodes corresponding to the semantic information of the voice information to be interacted, the number of nodes corresponding to the semantic information of the text information to be interacted, and the number of nodes corresponding to the semantic information of the video to be interacted are all one, the sum of the node distances between the nodes is obtained. When the sum of the node distances between the nodes is less than a first preset threshold, it is determined that the obtained node distances between the nodes determine that the obtained voice information to be interacted, the video information to be interacted, and the text information to be interacted need to be transmitted to the client at the same time.
8. The digital human interaction method based on a large language model as claimed in claim 7, characterized in that: The client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, further comprising: When the number of nodes corresponding to the semantic information of the voice information to be interacted is one and the number of nodes corresponding to the semantic information of the text information to be interacted and the number of nodes corresponding to the semantic information of the video information to be interacted are more than one, the following method is used to determine whether to transmit them to the client at the same time: Obtaining the node distances between the nodes corresponding to the semantic information of each speech information to be interacted and the nodes of the semantic information of each text information to be interacted, the node distances being referred to as first node distances; Obtaining the node distances between the nodes corresponding to the semantic information of each to-be-interacted voice information and the nodes corresponding to the semantic information of each to-be-interacted video information, the node distances being referred to as second node distances; When the sum of a first node distance and a second node distance is less than a first preset threshold, the node distances between the acquired nodes are determined to determine whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
9. The digital human interaction method based on a large language model as claimed in claim 8, characterized in that: The client determines whether the acquired voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time according to the acquired node distances between the nodes, further comprising: When the number of nodes corresponding to the semantic information of the voice information to be interacted, the number of nodes corresponding to the semantic information of the text information to be interacted, and the number of nodes corresponding to the semantic information of the video information to be interacted are all multiple, the following method is used to determine whether to transmit them to the client at the same time: A scatter plot is made with any node as the origin. All nodes are located on the scatter plot, and the shortest straight line distance between each node and the origin is the node distance between the node and the origin. Cluster each node on the scatter plot to obtain clusters; When each node in a cluster has at least one semantic information belonging to the voice information to be interacted, at least one node corresponding to the semantic information belonging to the text information to be interacted, and at least one node belonging to the semantic information of the video to be interacted, then the node distance between the obtained nodes is determined to determine whether the obtained voice information to be interacted, video information to be interacted, and text information to be interacted need to be transmitted to the client at the same time.
10. A digital human interaction device based on a large language model, characterized in that: The client and cloud of the digital human interaction device based on the large language model cooperate to implement the digital human interaction device based on the large language model as described in any one of claims 1 to 9.