A target person's speech data extraction method, system, device and storage medium
Patent Information
- Application Number
- CN202210253016.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-03-15
Smart Images

Figure CN114863930B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a target person voice data extraction method, system, device and storage medium. BACKGROUND
[0002] In the field of public security, it is necessary to effectively supervise certain target persons and collect their relevant information for analysis and investigation.
[0003] In related technologies, it is difficult to find suitable control points in the supervision process, and a large amount of manpower and material resources need to be invested in the design of control points. The traditional method is to use face recognition technology to collect face information at the possible locations of the target person to determine their possible whereabouts. This method has a lot of limitations: for example, the image accuracy collected by the face recognition technology is not high in the case of poor angle and dark light, and the face recognition effect is not good, resulting in low accuracy and poor usability of the information.
[0004] In summary, the problems in related technologies need to be solved urgently. SUMMARY
[0005] The present application aims to at least partially solve one of the technical problems in the prior art.
[0006] To this end, one object of the present application is to provide a target person voice data extraction method, system, device and medium, which can improve the accuracy and usability of the collected target person information.
[0007] In order to achieve the above technical purpose, the technical solution adopted by the embodiments of the present application comprises:
[0008] On the one hand, the present application provides a target person voice data extraction method, comprising the following steps:
[0009] Obtaining the voice data of a person collected by a target device installed at a predetermined position;
[0010] Matching the voice data of the person with a voiceprint database corresponding to the target person to determine whether the voice data of the person includes target voice data;
[0011] When the voice data of the person includes target voice data, extracting and saving the target voice data;
[0012] Wherein, the installation position of the target device is determined by the following steps:
[0013] Establishing a knowledge graph of the target person;
[0014] According to the knowledge graph, prompt information is output, and the prompt information is used to instruct a user to install the target device at a specified position.
[0015] Further, after the target voice data is extracted and saved, the method further includes the following steps:
[0016] Voice recognition is performed on the target voice data to obtain text content of the target voice data.
[0017] Text feature information of the text content is extracted.
[0018] The text feature information is input into a prediction model to obtain a behavior prediction result of the target person.
[0019] Further, the method further includes the following steps:
[0020] The target voice data is input into a noise detection model to obtain a noise detection result output by the noise detection model; the noise detection result is used to represent whether the target voice data contains noise data.
[0021] According to the noise detection result, a confidence degree of the behavior prediction result is determined.
[0022] Further, the method further includes the following steps:
[0023] Image data collected by a target device installed at the preset position is obtained.
[0024] Face recognition features in the image data are extracted.
[0025] The face recognition features are matched with a face database corresponding to the target person to determine whether the image data includes the target person.
[0026] Further, the step of obtaining the person voice data collected by the target device installed at the preset position includes:
[0027] Raw voice data collected by the target device installed at the preset position is obtained.
[0028] Voiceprint features and personal features corresponding to all persons in the raw voice data are extracted.
[0029] According to the voiceprint features, a person voice model corresponding to the personal features is constructed.
[0030] The raw voice data is processed through the person voice model to obtain the person voice data.
[0031] Further, the step of matching the person voice data with the voiceprint database corresponding to the target person to determine whether the target voice data is included in the person voice data comprises:
[0032] extracting first voiceprint features of the person voice data;
[0033] extracting second voiceprint features corresponding to the target voice data of the target person from the voiceprint database;
[0034] determining similarity of the first voiceprint features and the second voiceprint features;
[0035] determining whether the similarity is greater than a preset threshold;
[0036] when the similarity is greater than the preset threshold, the target voice data is included in the person voice data;
[0037] when the similarity is less than or equal to the preset threshold, the target voice data is not included in the person voice data.
[0038] Further, the step of establishing the knowledge graph of the target person comprises:
[0039] obtaining personal information of the target person, the personal information including family information and social relationship of the target person;
[0040] determining main activity places of the target person and associated persons according to the personal information of the target person;
[0041] constructing the knowledge graph according to the personal information and the main activity places.
[0042] On the other hand, the embodiment of the present application proposes a voice data extraction system of a target person, comprising:
[0043] a first module for obtaining person voice data collected by a target device installed at a preset position;
[0044] a second module for matching the person voice data with a voiceprint database corresponding to the target person to determine whether the target voice data is included in the person voice data;
[0045] a third module for extracting and saving the target voice data when the target voice data is included in the person voice data;
[0046] wherein, the installation position of the target device is determined by the following module:
[0047] a fourth module for establishing a knowledge graph of the target person;
[0048] a fifth module configured to output prompt information according to the knowledge graph, the prompt information being used to instruct a user to install the target device at a specified position.
[0049] In another aspect, an embodiment of the present application provides a device for extracting voice data of a target person, comprising:
[0050] at least one processor;
[0051] at least one memory configured to store at least one program;
[0052] The at least one program, when executed by the at least one processor, causes the at least one processor to implement the method for extracting voice data of a target person.
[0053] In another aspect, an embodiment of the present application provides a storage medium having processor-executable instructions stored therein, the processor-executable instructions, when executed by a processor, being used to implement the method for extracting voice data of a target person.
[0054] The present application discloses a method for extracting voice data of a target person, which has the following advantages:
[0055] The present embodiment establishes a knowledge graph of a target person, and outputs prompt information according to the knowledge graph to instruct a user to install a target device at a specified position. Then, the voice data of a person collected by the target device installed at the preset position is acquired, and the voice data of the person is matched with a voiceprint database corresponding to the target person to determine whether the target voice data is included in the voice data of the person. When the target voice data is included in the voice data of the person, the target voice data is extracted and saved. This method can predict a more appropriate control point through a knowledge graph in a supervision process, reduce the large investment of manpower and material resources, and match the acquired voice data of a person through a target person voiceprint database, which is conducive to outputting a more accurate matching result to improve the accuracy and usability of target person information. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the drawings in the following introduction are only used to facilitate the clear description of some embodiments in the technical solutions of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0057] Figure 1 A module schematic diagram of a remote monitoring platform provided by an embodiment of the present application is shown in the following figure:
[0058] Figure 2 A hardware structure schematic diagram of a remote monitoring platform provided for an embodiment of the present application is shown in the figure.
[0059] Figure 3 A flowchart of a target person's voice data extraction method provided for an embodiment of the present application is shown in the figure.
[0060] Figure 4 A flowchart of a target device installation position determination method provided for an embodiment of the present application is shown in the figure.
[0061] Figure 5 A flowchart of a human voice separation method provided for an embodiment of the present application is shown in the figure.
[0062] Figure 6 A flowchart of a person voice matching method provided for an embodiment of the present application is shown in the figure.
[0063] Figure 7 A flowchart of a person relationship knowledge graph determination method provided for an embodiment of the present application is shown in the figure.
[0064] Figure 8 A flowchart of a target person's behavior prediction method provided for an embodiment of the present application is shown in the figure.
[0065] Figure 9 A flowchart of a face recognition method provided for an embodiment of the present application is shown in the figure.
[0066] Figure 10 A structure schematic diagram of a target person's voice data extraction system provided for an embodiment of the present application is shown in the figure.
[0067] Figure 11 A structure schematic diagram of a target person's voice data extraction device provided for an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0068] This part will describe the specific embodiments of the present application in detail, the preferred embodiments of the present application are shown in the attached drawings, the role of the drawings is to supplement the description of the text part with figures, so that people can intuitively and visually understand each technical feature and the overall technical scheme of the present application, but it cannot be understood as a limitation on the protection scope of the present application.
[0069] In the description of the embodiments of the present application, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, more than, etc. are understood as not including the number, above, below, within, etc. are understood as including the number, "at least one" means one or more, "at least one of the following" and the like means any combination of the items, including any combination of single or multiple items. If there is a description of "first", "second", etc. is only used to distinguish technical features for the purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or implicitly indicating the order of the indicated technical features.
[0070] It should be noted that the terms such as setting, installing, connecting, etc. in the embodiments of the present application should be understood broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in the embodiments of the present application in combination with the specific content of the technical solutions. For example, the term "connection" can be mechanical connection, or electrical connection or can communicate with each other; it can be directly connected, or indirectly connected through an intermediate medium.
[0071] In the description of the embodiments of the present application, the description of the terms "one embodiment", "another embodiment" or "some embodiments", "in the above embodiment" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least two embodiments or embodiments of the present disclosure. In the present disclosure, the illustrative description of the above terms does not necessarily refer to the same example or embodiment. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or embodiments in a suitable manner.
[0072] It should be noted that the technical features involved in each of the embodiments of the present application described below can be combined with each other as long as there is no conflict.
[0073] First, the terms involved in the present application are analyzed:
[0074] Artificial intelligence (AI): is a new technical science of studying, developing the theory, method, technology and application system for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science, artificial intelligence attempts to understand the essence of intelligence, and produce a new intelligent machine that can react in a similar way to human intelligence, the research in this field includes robots, language recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0075] Natural language processing (NLP): NLP uses computers to process, understand and use human language (such as Chinese, English, etc.), NLP is a branch of artificial intelligence, and is a cross-discipline of computer science and linguistics, and is also commonly known as computational linguistics. Natural language processing includes syntax analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information retrieval, information extraction and filtering, text classification and clustering, public opinion analysis and opinion mining, etc. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and language computing related linguistic research.
[0076] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0077] Artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc. several directions.
[0078] In order to protect public safety, it is necessary to monitor some target persons. At present, the monitoring method of the target person mainly includes the following steps: arranging control at the possible location of the target person, collecting the face information of the target person by using the image device such as camera, and comparing the face information by using the face recognition technology, so as to realize the monitoring of the target person. However, this monitoring method has the following problems: on the one hand, the possible location of the target person is uncertain, which leads to the need of a large amount of manpower and material resources for arrangement; on the other hand, the face recognition technology has the problems of complex and diverse background interference, complex and changeable illumination conditions, different face posture and expression changes, and external occlusion, which leads to the fact that the face information cannot be correctly extracted and the face recognition rate is low.
[0079] Therefore, the present application provides a voice data extraction method, system and device of a target person and a storage medium. The method includes the following steps: establishing a knowledge graph of the target person, and outputting prompt information according to the knowledge graph to instruct a user to install a target device at a specified location. Then, the voice data of a person collected by the target device installed at the preset location is acquired, and the voice data of the person is matched with a voiceprint database corresponding to the target person to determine whether the target voice data is included in the voice data of the person. When the target voice data is included in the voice data of the person, the target voice data is extracted and saved. This method can predict a more appropriate control point in the monitoring process by using the knowledge graph, reduce the large investment of manpower and material resources, and match the acquired voice data of the person with the voiceprint database of the target person, which is beneficial to output a more accurate matching result, so as to improve the accuracy and availability of the information of the target person.
[0080] Referring to Figure 1 In the embodiments of the present application, a remote monitoring platform is provided, which includes a data acquisition module 110, a control center 120, a database 130 and a display module 140. The data acquisition module 110 is used to acquire the voice data and image data of a person collected by a target device installed at a preset location, and send the data to the control center 120 for further processing. The database 130 stores the personal information, voice data and face data of the target person. The control center 120 is used to construct a corresponding knowledge graph according to the personal information of the target person stored in the database 130, and output prompt information for instructing a user to install a device at a specified location according to the knowledge graph, and display the prompt information by using the display module 140. The control center 120 can also compare the features in the voiceprint database and the face database of the database 130 with the voice data and image data of the person collected by the target device, so as to determine whether the voice data of the target person is included in the voice data of the person collected by the target device.
[0081] It should be noted that only a part of the modules of the remote monitoring platform is exemplarily given in the embodiments of the present application, and the remote monitoring platform can further include other relevant modules included in the remote monitoring platform to realize corresponding functions, and the specific implementation is not limited.
[0082] Referring to Figure 2 , Figure 2 is a schematic diagram of a hardware structure of the remote monitoring platform involved in the embodiments of the present application. In the embodiments of the present application, the remote monitoring platform can include a processor 210 (for example, a central processing unit, CPU), a communication bus 220, an input port 230, an output port 240, and a memory 250. The communication bus 220 is used to realize the connection and communication among the components; the input port 230 is used for data input; the output port 240 is used for data output, and the memory 250 can be a high-speed RAM memory or a stable memory (for example, a disk memory), and the memory 250 can optionally be a storage device independent of the aforementioned processor 210. Those skilled in the art can understand that Figure 2 The hardware structure shown in the foregoing embodiment does not constitute a limitation on the present application, and can include more or fewer components than those shown in the diagram, or combine certain components, or different component arrangements.
[0083] Continuing to refer to Figure 2 , Figure 2 The memory 250 as a readable storage medium in the foregoing embodiment can include an operating system, a network communication module, an application program module, and a control program. In the foregoing embodiment, Figure 2 , the network communication module is mainly used to connect a server and perform data communication with the server; and the processor 210 can call the control program of the remote monitoring platform stored in the memory 250 and execute the target person's voice data extraction method provided in the embodiments of the present application.
[0084] Based on the remote monitoring platform shown in Figure 1 and Figure 2 , as shown in Figure 3 , the embodiments of the present application provide a target person's voice data extraction method, which includes but is not limited to steps S310, S320, and S330:
[0085] S310, acquiring person voice data collected by a target device installed at a preset position;
[0086] In step S310, the character voice data mainly includes voice data collected at a preset position. Specifically, in the embodiment, the acquisition channel of the character voice data is not limited, and the character voice data can be collected from the preset position by a sound collecting device directly, or acquired from other electronic devices and computer systems through a data transmission interface or remote communication transmission. The target device installed at the preset position can be a sound collecting device, or an input and output device such as a smart voice interaction device, a smart phone and a tablet computer.
[0087] Referring to Figure 5 In some embodiments, step S310 can be further divided into steps S311-S314:
[0088] S311, acquiring original voice data collected by a target device installed at a preset position;
[0089] S312, extracting voiceprint features and personal features corresponding to all characters in the original voice data;
[0090] S313, constructing a voice model corresponding to the personal features according to the voiceprint features;
[0091] S314, processing the original voice data through the voice model to obtain the character voice data.
[0092] In the process of collecting character voice data, it is inevitable that the original voice data collected contains voice data of multiple characters. Based on this situation, the original voice data collected can be extracted according to different characters, and a corresponding voice model can be constructed for different characters, and then the voice data of the characters can be extracted from the original voice data through the voice model. Specifically, the voiceprint features and personal features of each speaking character can be collected, and a voice model associated with the personal features and containing the voiceprint features can be constructed. Then, the character voice data is separated from the original voice data through the constructed voice model, and each speaking character in the audio content after voice separation is labeled with the personal features in the form of time stamp, that is, different speaking characters can be separated in real time when listening to a whole recording.
[0093] Referring to Figure 4 In some embodiments, the installation position of the target device in step S310 can be determined by the following steps:
[0094] S410, establishing a knowledge graph of the target character.
[0095] S420, outputting prompt information according to the knowledge graph, the prompt information being used for instructing the user to install the target device at a specified position.
[0096] It can be understood that the big data statistical model can only find the correlation, and lacks deep logical analysis, and therefore it is necessary to construct a knowledge graph for further reasoning analysis. The knowledge graph includes a mode layer and a data layer, the mode layer is composed of concepts with hierarchical relationships, and the data layer is composed of entities of the concepts and relationships between the entities.
[0097] The knowledge graph of the embodiment of the application is a structured network composed of nodes and relationships, and records various dynamic relationships and static attributes of social relationships, family information and work conditions of the target person and associated persons. Through the knowledge graph, the dynamic relationships and static attributes are managed in a structured manner, the prompt information instructing the user to install the target device at a specified position is outputted, the usability of the information collected by the target device can be improved, the unstable factors of setting the target device artificially can be reduced, and therefore the consumption cost of manpower and material resources can be reduced.
[0098] Referring to Figure 7 In some embodiments, the step S410 can be further divided into steps S411-S413.
[0099] S411, obtaining personal information of the target person, the personal information including family information and social relationship of the target person;
[0100] S412, determining main activity places of the target person and associated persons according to the personal information of the target person;
[0101] S413, constructing a knowledge graph according to the personal information and the main activity places.
[0102] The key of constructing the knowledge graph is knowledge acquisition, and the main tasks of the knowledge acquisition include entity recognition, relationship extraction, attribute extraction, knowledge graph completion and other entity-oriented acquisition tasks. Among them, the relationship extraction is a key task of automatically constructing a large knowledge graph, which extracts unknown relationship facts from naive text and adds them to the knowledge graph.
[0103] In the embodiment of the present application, the construction of the knowledge graph adopts a "top-down" method, first models the knowledge points, concepts and terms possessed by the field, extracts the most extensive concepts, and then gradually refines them to define more attributes and relationships to constrain more specific categories. For example, for the public security field, first define the ontology concept "target person", and expand the "associated persons" related to the "target person" according to the "social relationship". The social relationship can be a first-degree social relationship, a second-degree social relationship, a third-degree social relationship and a fourth-degree social relationship. Different "associated persons" can be expanded according to different social relationships to form a person relationship knowledge graph of the target person. Through the structured data set and the knowledge base, the corresponding entity set of each concept is obtained, and on this basis, the field entities in semi-structured and unstructured data are automatically extracted through natural language processing related technologies, and the person relationship knowledge graph is automatically constructed through statistical, clustering, named entity recognition, relationship extraction, attribute extraction and other analysis methods. The specific process of automatically constructing the person relationship knowledge graph is as follows:
[0104] First, the family information and the first-degree person relationship of the target person need to be obtained, or the second-degree and third-degree person relationships of the associated persons are obtained as the seed entity set.
[0105] Then, according to the obtained personal information of the target person, the main activity places of the target person and the associated persons are determined, and the main activity places are associated with the seed entity set of the knowledge graph.
[0106] Next, the knowledge elements defined by the knowledge modeling are instantiated and obtained, including concept instantiation, relationship extraction and attribute extraction.
[0107] The instantiation of the entity concept needs to extract entities from the data. For structured data, high-frequency terms and high-frequency words are added to the entities of the corresponding concept; for unstructured data, entity recognition technology is used for modeling, and the extracted entities are used as instances of the corresponding concept.
[0108] After obtaining a series of discrete entity concepts, the relationships between the concepts need to be established through relationship extraction. Common association relationships include: isA (inheritance relationship), hasA (composition relationship), useA (dependence relationship) and other association relationships.
[0109] After establishing the relationship between the concepts, attribute information needs to be supplemented for each concept to fully describe the entity itself. On the one hand, the definition, annotation and explanation of the field terms are referred to to add attributes to the concept. On the other hand, attribute annotation models are established on structured data and unstructured data to extract and train the data, and then the entity attribute extraction is performed.
[0110] S320, match the person voice data with the voiceprint database corresponding to the target person to determine whether the target voice data is included in the person voice data;
[0111] In step S320, the obtained person voice data needs to be matched with the voice data saved in the target person voiceprint database to determine whether the target person's voice is included in the person voice data, so as to determine whether the target person has ever appeared at the preset location. First, generally, the person voice data is unstructured data, for the convenience of processing, the feature extraction is needed, and then the extracted voiceprint features are input into the corresponding machine learning model for comparison, and the approximation degree of the voiceprint features of the person voice data and the voiceprint features of the target person is output, so as to determine whether the target voice data is included in the person voice data.
[0112] Referring to Figure 6 In some embodiments, step S320 can be further divided into steps S321-S324:
[0113] S321, extract the first voiceprint features of the person voice data;
[0114] S322, extract the second voiceprint features corresponding to the voice data of the target person from the voiceprint database;
[0115] S323, determine the similarity of the first voiceprint features and the second voiceprint features;
[0116] S324, determine whether the similarity is greater than a preset threshold; when the similarity is greater than the preset threshold, the target voice data is included in the person voice data; when the similarity is less than or equal to the preset threshold, the target voice data is not included in the person voice data.
[0117] Specifically, the first voiceprint features here can include the acoustic feature information of the person voice data, which can be the digital features of the audio spectrum of the person voice data, for example. Specifically, in some embodiments, some time-frequency points can be selected from the audio spectrum of the person voice data according to a predetermined rule, and encoded into a digital sequence, which can be used as the acoustic feature information of the person voice data. For example, the acoustic feature information can also be extracted based on the dimensions of pronunciation accuracy, fluency, prosody, signal-to-noise ratio, sound intensity, etc. in the present application. And in some embodiments, the acoustic feature information extracted from multiple dimensions can also be integrated to obtain new acoustic feature information.
[0118] The second voiceprint feature can include acoustic feature information of the voice data of the target person, and the second voiceprint feature can be extracted from the voice data of the target person. In this embodiment, the channel for obtaining the voice data of the target person is not limited, and the voice data of the target person can be collected from a preset place through a sound collecting device, or can be obtained from other devices through remote communication transmission. Since the voice data of the person is unstructured data, for the convenience of processing, the feature information of the voice data of the person is extracted in this application, and the extracted feature information is recorded as the first voiceprint feature, and the feature information of the voice data of the target person stored in the voiceprint database is recorded as the second voiceprint feature.
[0119] In the matching process, a targeted machine learning model can be trained to calculate the similarity, and the similarity of the first voiceprint feature and the second voiceprint feature is output. The similarity here is used to represent the degree of similarity between the first voiceprint feature and the second voiceprint feature. When the value of the similarity reaches a certain value, it can be considered that the first voiceprint feature and the second voiceprint feature are the same, and it can be considered that the voice data of the person includes the target voice data. In addition, in some embodiments, a vector index can be set for the acoustic feature information in the form of a vector to reduce the data operation amount in the matching query process.
[0120] Specifically, when determining the similarity between the first voiceprint feature and the second voiceprint feature, in some embodiments, the difference value between the first feature information and the second feature information can be determined first, and then the similarity can be determined according to the difference value. The greater the difference value, the smaller the similarity, and vice versa, the smaller the difference value, the greater the similarity.
[0121] It can be understood that in actual application, even if two pieces of voice data are emitted by the same person, the similarity is not necessarily 100%, which is caused by environmental factors and incomplete training of the model. Therefore, a preset threshold value needs to be set, which can also be understood as a matching error range. When the similarity is greater than the preset threshold value, it is considered that the voice data of the person includes the target voice data within the error allowable range; when the similarity is less than or equal to the preset threshold value, it is considered that the voice data of the person does not include the target voice data. The preset threshold value can be set by those skilled in the art according to actual conditions, which is not limited here.
[0122] S330, when the voice data of the person includes the target voice data, the target voice data is extracted and saved.
[0123] At step S330, according to the foregoing description, when the similarity between the first voiceprint feature and the second voiceprint feature is greater than a preset threshold, it is considered that the target voice data is included in the person voice data. For the field of public security, in order to facilitate further control and reconnaissance, the target voice data can be extracted and saved. In some embodiments, the sound segment of the target person detected can be intercepted to facilitate manual checking by the user.
[0124] With reference to Figure 8 In the embodiments of the present application, a behavior prediction method of a target person is also provided, which mainly includes steps S510 to S550.
[0125] S510, performing voice recognition on the target voice data to obtain text content of the target voice data;
[0126] S520, extracting text feature information of the text content;
[0127] S530, inputting the text feature information into a prediction model to obtain a behavior prediction result of the target person;
[0128] S540, inputting the target voice data into a noise detection model to obtain a noise detection result output by the noise detection model; the noise detection result is used to represent whether the target voice data contains noise data;
[0129] S550, determining a confidence degree of the behavior prediction result according to the noise detection result.
[0130] In the embodiments of the present application, when extracting the text feature information, the target voice data needs to be textually processed first. The automatic speech recognition technology (ASR) can be used to perform voice recognition on the target voice data to obtain the text content of the target voice data, and then the text feature information of the text content is extracted. For example, the text content of the target voice data can be converted into structured data such as a vector through natural language processing technology, so as to take the converted structured data as the text feature information. Then, the text feature information can be input into a pre-trained prediction model to obtain the behavior prediction result of the target person. Specifically, the form of the behavior prediction result can be flexibly set according to needs, and a prediction model is built by selecting a suitable algorithm. For example, the prediction model can extract the text content based on the target voice data, and then perform semantic analysis according to the extracted text feature information, so as to predict the behavior action that the target person may take.
[0131] It can be understood that the target person behavior prediction result obtained by the prediction model is not necessarily completely reliable, and therefore completely relying on the aforementioned prediction model to predict the behavior of the target person may result in prediction errors. Therefore, in the embodiments of the present application, noise detection analysis is performed on the target speech data based on the text feature information to assist in judging the reliability of the behavior prediction result. The noise detection analysis is used to analyze whether the target speech data contains noise. The above-mentioned situation can be analyzed by modeling and training a targeted machine learning model to analyze noise, and output whether there is noise or the degree of influence of noise. For example, in the embodiments of the present application, a noise detection model can be used to detect whether the target speech data contains noise data. Specifically, at this time, the aforementioned target speech data can be input into the noise detection model, and the noise detection model processes the target speech data and outputs a noise detection result. Of course, in the embodiments of the present application, the noise detection model can be further subdivided, for example, an environmental noise model can be established to detect environmental noise in the target speech data, a human voice noise model can be established to detect human voice noise in the target speech data, and the like. After the noise detection result of the target speech data is determined, the reliability of the behavior prediction result obtained from the target speech data can be effectively quantified, that is, the confidence of the behavior prediction result.
[0132] Referring to Figure 9 In some embodiments, in order to further improve the accuracy of collecting information of the target person, the embodiments of the present application also provide a face recognition method, which can also be applied to the supervision process of the target person. The method obtains image data of the target person from another dimension different from speech data, which is conducive to the identification of the identity of the target person. The method mainly includes steps S610 to S630:
[0133] S610, obtaining image data collected by a target device installed at the preset position;
[0134] S620, extracting face recognition features in the image data;
[0135] S630, matching the face recognition features with a face database corresponding to the target person to determine whether the image data includes the target person.
[0136] In the embodiment, the face recognition method can be used to assist in judging the identity of the target person. When the similarity meets a certain range, the face recognition technology can be called to further confirm the identity of the target person. For example, the preset similarity threshold is set to 95%, but the actual output similarity is 92%, which is less than 95%. However, the similarity is still in a high range, and the machine learning model and environmental factors can be too large to cause the similarity to not meet the threshold. Therefore, when the similarity is in the range of 90%-94%, the face recognition technology can be called to further confirm the identity of the suspect to correct the similarity value. In addition, the voiceprint recognition technology and the face recognition technology can be used simultaneously to output the voiceprint similarity and the face similarity, and then the voiceprint similarity and the face similarity are weighted and summed to obtain the final similarity.
[0137] In the embodiment, the image data collected by the target device installed at the preset position is needed. The portrait data mainly includes the face image data collected at the preset position. The channel for obtaining the voice data of the person is not limited, and the voice data of the person can be collected from the preset position by a sound collecting device, or obtained from other electronic devices and computer systems through a data transmission interface or remote communication transmission. The installation position of the target device in the embodiment can also be determined by steps S410-S420, which will not be described here. After obtaining the image data, the image data needs to be feature extracted to obtain the face recognition features. Then, the face recognition features are matched with the target recognition features in the face database of the target person to determine whether the image information collected includes the face image information of the target person, so as to confirm whether the target person appears at the preset position. The above method can be realized by a pre-trained convolutional neural network. Specifically, the data in the training data set with labels is input into the initialized convolutional neural network, and the recognition result output by the model, i.e., the predicted recognition result, can be obtained. The accuracy of the recognition model can be evaluated according to the predicted recognition result and the aforementioned label, so that the parameters of the model are updated.
[0138] For a face recognition model, the accuracy of the model recognition result can be measured by a loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of a training data is determined by the label of the training data and the prediction result of the training data by the model. In actual training, a training data set has many training data. Therefore, a cost function is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction error of all training data, which can better measure the prediction effect of the model. For a general machine learning model, based on the aforementioned cost function, plus a regular term that measures the complexity of the model, the target function for training can be obtained. Based on the target function, the loss value of the entire training data set can be obtained. There are many commonly used loss functions, such as 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross-entropy loss function, etc. All of them can be used as the loss function of the machine learning model. Here, they will not be elaborated one by one. In the embodiment of the present application, any one of the loss functions can be selected to determine the loss value of the training. Based on the loss value of the training, the parameters of the model are updated by using the back propagation algorithm, and after several iterations, a trained face recognition model can be obtained. The specific number of iterations can be pre-set, or the training is considered to be completed when the test set reaches the accuracy requirement.
[0139] In some embodiments, a display method of the matching result is also provided. The display method can be applied to a terminal device, for example, can be applied to part of the software in the terminal device, and is used to realize part of the software function. Similarly, the terminal device to which the display method can be applied includes but is not limited to a smart watch, a smart phone, a tablet computer, a personal digital assistant (PDA), a smart voice interaction device, a notebook computer, a desktop computer, a smart home appliance, or a vehicle-mounted terminal.
[0140] In some embodiments, the display manner of the matching result can be displaying the similarity of the matching result, or whether the target person corresponding voice data is matched, the display manner of the matching result can be directly reminding in the form of text in the display screen of the terminal such as the touch display screen, the smart phone or the display interface of the APP, the text can be Chinese characters or characters of other countries. Alternatively, the display manner of the matching result can also be switching the display color of the preset display area of the matching result from the first color (such as green) to the second color (such as red and the like) in the display screen of the terminal such as the touch display screen, the smart phone or the display interface of the APP. In some embodiments, an alarm system can also be set, when the matching result considers that the voice data of the person includes the target voice data, that is, the similarity of the matching result is higher than the preset threshold, an alarm information is generated to remind the user to pay attention to the matching result.
[0141] From the above, the application establishes the knowledge graph of the target person, and outputs the prompt information according to the knowledge graph to instruct the user to install the target device at the specified position. Then, the voice data of the person collected by the target device installed at the preset position is acquired, and the voice data of the person is matched through the voiceprint database corresponding to the target person to determine whether the target voice data is included in the voice data of the person, and when the voice data of the person includes the target voice data, the target voice data is extracted and saved. This method can predict a more appropriate control point in the supervision process through the knowledge graph, reduce the large investment of manpower and material resources, and match the acquired voice data of the person through the target person voiceprint database, which is conducive to outputting a more accurate matching result to improve the accuracy and usability of the target person information.
[0142] Referring to Figure 10 The embodiment of the application provides a voice data extraction system of a target person, which comprises:
[0143] A first module 1001 is configured to acquire voice data of a person collected by a target device installed at a preset position;
[0144] A second module 1002 is configured to match the voice data of the person through a voiceprint database corresponding to the target person to determine whether the target voice data is included in the voice data of the person;
[0145] A third module 1003 is configured to extract and save the target voice data when the voice data of the person includes the target voice data;
[0146] The installation position of the target device is determined by the following module:
[0147] A fourth module 1004 is configured to establish a knowledge graph of the target person;
[0148] A fifth module 1005 is configured to output prompt information according to the knowledge graph, the prompt information being used to instruct a user to install a target device at a specified position.
[0149] The content in the method embodiments is applicable to the system embodiments, the system embodiments achieve the same functions as the method embodiments, and achieve the same beneficial effects as the method embodiments.
[0150] Referring to Figure 11 The embodiment of the present application provides a target person voice data extraction device, which comprises:
[0151] At least one processor 1101;
[0152] At least one memory 1102 is configured to store at least one program;
[0153] When the at least one program is executed by the at least one processor 1101, the at least one processor 1101 is caused to implement Figure 3 The target person voice data extraction method shown in the figure.
[0154] The content in the method embodiments is applicable to the device embodiments, the device embodiments achieve the same functions as the method embodiments, and achieve the same beneficial effects as the method embodiments.
[0155] The embodiment of the present application further provides a storage medium, wherein the storage medium stores processor-executable instructions, and the processor-executable instructions are used to implement Figure 3 The target person voice data extraction method shown in the figure.
[0156] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the above-mentioned embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.
Claims
1. A voice data extraction method for a target person, characterized by, The method comprises the following steps: acquiring character voice data collected by a target device installed at a preset position; matching the character voice data with a voiceprint database corresponding to a target character to determine whether the character voice data includes target voice data; when the character voice data includes target voice data, extracting and saving the target voice data; performing voice recognition on the target voice data to obtain text content of the target voice data; extracting text feature information of the text content; inputting the text feature information into a prediction model to obtain a behavior prediction result of the target character, wherein the behavior prediction result is used to represent a predicted behavior action of the target character; inputting the target voice data into a noise detection model to obtain a noise detection result output by the noise detection model, wherein the noise detection result is used to represent whether the target voice data includes noise data; determining a confidence degree of the behavior prediction result according to the noise detection result; wherein the installation position of the target device is determined by the following steps: establishing a knowledge graph of the target character; outputting prompt information according to the knowledge graph, wherein the prompt information is used to instruct a user to install the target device at a specified position.
2. The voice data extraction method of a target person according to claim 1, characterized by, The method further comprises the following steps: acquiring image data collected by the target device installed at the preset position; extracting face recognition features in the image data; matching the face recognition features with a face database corresponding to the target character to determine whether the image data includes the target character.
3. The voice data extraction method of a target person according to claim 1, characterized by, The step of acquiring character voice data collected by a target device installed at a preset position comprises: acquiring original voice data collected by the target device installed at the preset position; extracting voiceprint features and personal features corresponding to all characters in the original voice data; constructing a voice model corresponding to the personal features according to the voiceprint features; processing the original voice data through the voice model to obtain the character voice data.
4. The voice data extraction method of a target person according to claim 1, characterized by, The step of matching the character voice data with a voiceprint database corresponding to a target character to determine whether the character voice data includes target voice data comprises: extracting first voiceprint features of the character voice data; extracting second voiceprint features corresponding to voice data of the target character from the voiceprint database; determining a similarity between the first voiceprint features and the second voiceprint features; determining whether the similarity is greater than a preset threshold; when the similarity is greater than the preset threshold, the character voice data includes target voice data; when the similarity is less than or equal to the preset threshold, the character voice data does not include target voice data.
5. The voice data extraction method of a target person according to claim 1, characterized by, The step of establishing a knowledge graph of the target character comprises: acquiring personal information of the target character, wherein the personal information includes family information and social character relationships of the target character; determining main activity places of the target character and associated characters according to the personal information of the target character; According to the personal information and the main activity place, the knowledge graph is constructed.
6. A voice data extraction system for a target person, characterized by comprising: The method comprises the steps of: A first module is configured to acquire character voice data collected by a target device installed at a preset position; A second module is configured to match the character voice data with a voiceprint database corresponding to a target character to determine whether the character voice data includes target voice data; A third module is configured to extract and save the target voice data when the character voice data includes the target voice data; Voice recognition is performed on the target voice data to obtain text content of the target voice data; Text feature information of the text content is extracted; The text feature information is input into a prediction model to obtain a behavior prediction result of the target character, wherein the behavior prediction result is used to represent a predicted behavior action of the target character; The target voice data is input into a noise detection model to obtain a noise detection result output by the noise detection model, wherein the noise detection result is used to represent whether the target voice data includes noise data; According to the noise detection result, a confidence degree of the behavior prediction result is determined; The installation position of the target device is determined by the following modules: A fourth module is configured to establish a knowledge graph of the target character; A fifth module is configured to output prompt information according to the knowledge graph, wherein the prompt information is used to instruct a user to install the target device at a specified position.
7. A device for extracting voice data of a target person, characterized in that, The method comprises the steps of: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the target character voice data extraction method according to any one of claims 1-5.
8. A computer-readable storage medium having stored therein instructions that are executable by a processor, the instructions comprising: The instructions executable by the processor are used to implement the target character voice data extraction method according to any one of claims 1-5 when executed by the processor.
Citation Information
Patent Citations
Anomaly detection method and device, computer storage medium and electronic equipment
CN110209835A
Medical information feedback method and device, equipment and readable storage medium
CN111813946A
Voice processing method and apparatus, system and storage medium
WO2022007497A1