Visual verification method and system
By collecting and processing transcripts and related texts, using semantic features for comparative analysis, generating intent-related result sets and factor sets, and finally displaying visual evidence content through an intent recognition prediction model, the problem of the single form of visual evidence content display in existing technologies is solved, and the deep integration and clear display of case information is achieved.
Patent Information
- Application Number
- CN202510688640.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing technologies have limitations when processing transcripts and related information. Traditional visual evidence content display is single in form, lacks intuitiveness and relevance, and it is difficult to quickly grasp the logical relationship between key information and evidence.
A visual evidence method is provided. By collecting transcript text and related associated text and background text, the text feature set and paragraph set are processed, and semantic features are used for comparison and analysis to generate intention association result set and intention factor set. Finally, the visual evidence content is displayed through the intention recognition prediction model.
It achieves an in-depth integration and display of transcripts and evidence materials, clearly presenting the overall picture of the case and the chain of evidence, and improving the analysis efficiency and accuracy of case handlers.
Smart Images

Figure CN120671673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visualization technology, and more particularly, to a visualization evidence method and system. Background Art
[0002] In related fields, transcripts are important information carriers, and accurate identification of their intent and effective presentation of visual evidence are crucial. Current technologies have certain limitations when processing transcripts and related information. Traditional visual evidence presentation is single-format, mostly static text lists or simple evidence stacking, lacking intuitiveness and relevance. In the visual evidence process, a large amount of text materials and scattered evidence are presented, making it difficult to quickly grasp the key information and the logical relationship between evidence. At the same time, the lack of visual assistance makes it difficult to deeply integrate the transcript content with the evidence materials, making it impossible to clearly present the overall picture of the case and the chain of evidence, which is not conducive to understanding the facts of the case. Summary of the Invention
[0003] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a visual verification method and system.
[0004] To achieve the above object, the present invention provides the following technical solutions: A visual verification method, comprising the following steps: Collecting the transcript text and associated text and background text associated with the transcript text; Processing the transcript text to obtain a transcript text feature set and an associated paragraph set, segmenting the associated text into paragraphs based on semantic features to obtain an associated text feature set, and segmenting the transcript text into paragraphs based on background text to obtain a background text feature set; Compare and analyze the associated text feature set and the associated paragraph set to obtain the intention association result set; After determining whether the transcript text feature set and the background text feature set are related to each other under the condition of subject relevance, the first intention factor set and the second intention factor set are obtained respectively; After inputting the first intention factor set and the second intention factor set into the intention recognition prediction model, the intention recognition prediction result is obtained. The intention recognition prediction result is combined with the intention association result set to identify the intention of the transcript text and then display the visual evidence content.
[0005] Preferably, determining whether the transcript text feature set and the background text feature set are related to each other under the condition of subject relevance and performing processing and analysis to obtain a first intention factor set and a second intention factor set respectively, specifically includes the following steps: When the transcript text feature set and the background text feature set are in a topic-related condition, the transcript text feature set and the background text feature set are compared and analyzed to obtain the first intention factor set; When the transcript text feature set and the background text feature set are in the condition of being irrelevant to the subject, the associated text feature set and the background text feature set are compared and analyzed to obtain the second intention factor set.
[0006] Preferably, the transcript text is processed to obtain a transcript text feature set and an associated paragraph set, the associated text is segmented into paragraphs according to semantic features to obtain an associated text feature set, and the transcript text is segmented into paragraphs according to background text to obtain a background text feature set, specifically comprising the following steps: Extracting semantic distribution features of the transcript text, and segmenting the transcript text into several paragraphs based on the semantic distribution features to obtain a transcript paragraph set; The transcript paragraph set and the background text are segmented into several paragraphs according to their own semantic features to obtain the associated paragraph set and the background paragraph set; After determining the tags of the transcript paragraph set and the associated text according to the semantic importance, the transcript text feature set and the associated text feature set are obtained respectively; The background text feature set is obtained by marking the core intent paragraphs according to their respective semantic importance.
[0007] Preferably, comparing and analyzing the associated text feature set with the associated paragraph set to obtain an intention association result set specifically includes the following steps: Statistically associate the semantic keyword frequency dataset of each paragraph in the text feature set; Comparing the corresponding semantic keyword coincidence frequencies of the associated paragraph set and the paragraphs in the associated text feature set to obtain a first semantic coincidence frequency dataset; Calculating the proportion of the first semantic coincidence frequency dataset and the semantic keyword frequency dataset to obtain a semantic proportion dataset; When the semantic proportion data set contains a value greater than the preset semantic threshold, the difference exceeding the preset semantic threshold is calculated to obtain the semantic excess data set; After extracting the first intention information of the associated paragraph set, a scene context dataset is set; The intent association result set is obtained by inputting the first intent information, scene context dataset, associated text feature set and semantic excess dataset into the intent association value prediction model.
[0008] Preferably, when the transcript text feature set and the background text feature set are in a topic-related condition, the transcript text feature set and the background text feature set are compared and analyzed to obtain a first intention factor set, which specifically includes the following steps: When the transcript text feature set and the background text feature set are under a condition of being related to the subject, the transcript text feature set and the background text feature set are processed to obtain a second semantic coincidence frequency data set and a semantic sentiment tendency value; Counting the overlapping topics of each paragraph in the transcript text feature set and the background text feature set to obtain the overlapping topic dataset; The second intention information of the background text feature set is extracted, and the second semantic overlap frequency data set, semantic sentiment tendency value, overlap topic data set, second intention information, scene context data set, intention association result set and transcript text feature set are combined into a first intention factor set.
[0009] Preferably, when the transcript text feature set and the background text feature set are under a condition of being related to a topic, the transcript text feature set and the background text feature set are processed to obtain a second semantic overlap frequency dataset and a semantic sentiment tendency value, specifically comprising the following steps: When the transcript text feature set and the background text feature set are in a topic-related condition, the second semantic coincidence frequency data set is obtained by counting the frequency of coincidence of corresponding semantic keywords in each paragraph of the transcript text feature set and the background text feature set; Extract the semantic sentiment tendency value of the corresponding paragraph of the second semantic coincidence frequency dataset.
[0010] Preferably, when the transcript text feature set and the background text feature set are in a condition of being irrelevant to the subject, the associated text feature set and the background text feature set are compared and analyzed to obtain a second intention factor set, which specifically includes the following steps: When the transcript text feature set and the background text feature set are in a topic-unrelated condition, the topic information dataset is obtained by counting the topics between the corresponding paragraphs in the transcript text feature set and the background text feature set; Among them, the second semantic overlap frequency dataset, semantic sentiment tendency value, overlap topic dataset, second intention information, scene context dataset, transcript text feature set, intention association result set and topic information dataset are combined into a second intention factor set.
[0011] A visual evidence system, comprising: Collection module: collects the transcript text and the associated text and background text associated with the transcript text; Segmentation module: processes the transcript text to obtain a transcript text feature set and an associated paragraph set, segments the associated text into paragraphs based on semantic features to obtain an associated text feature set, and segments the transcript text into paragraphs based on background text to obtain a background text feature set; The first analysis module compares and analyzes the associated text feature set and the associated paragraph set to obtain the intention association result set; The second analysis module determines whether the transcript text feature set and the background text feature set are related to each other and processes and analyzes them to obtain the first intention factor set and the second intention factor set respectively; Display module: After inputting the first intention factor set and the second intention factor set into the intention recognition prediction model, the intention recognition prediction result is obtained. The intention recognition prediction result is combined with the intention association result set to identify the intention of the transcript text and then the visual evidence content is displayed.
[0012] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements a visual verification method when executing the program.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention mines the semantic connection between the associated text feature set and the associated paragraph set by statistically analyzing the frequency of semantic keywords and calculating the frequency of semantic overlap. In drunk driving cases, the degree of consistency between the statements of the parties and related texts such as laws and regulations, relevant cases, etc. is accurately judged. For example, the consistency of the parties' cognition and expression of key concepts such as "drunk driving" and "alcohol content standards" is accurately measured, which provides a strong basis for accurately determining whether they are aware that their behavior is illegal and reduces misjudgments caused by vague semantic understanding. When generating the first intention factor set and the second intention factor set, multi-dimensional information such as semantic overlap frequency, semantic sentiment tendency value, overlap topic data, intention information, scene context, etc. is integrated.
[0014] From extracting semantic distribution features from transcripts, segmenting paragraphs, to comparing and analyzing related text and paragraphs, this system automates and standardizes text processing. It rapidly filters, categorizes, and extracts features from large amounts of text. For example, when processing multiple drunk driving case transcripts, it can quickly generate text feature sets and paragraph sets, significantly reducing initial text processing time and improving case handling efficiency.
[0015] By tagging text according to semantic importance, the system can quickly locate key information within transcripts, related text, and background text. In drunk driving cases, this allows researchers to quickly identify core information such as "drinking amount," "driving time," and "alcohol test value," eliminating the need for blind searching within large amounts of text. This allows investigators to focus on analyzing key information and accelerate the case analysis process.
[0016] This technical solution comprehensively encompasses transcript text, related text (laws and regulations, similar cases, etc.), and background text (law enforcement records, environmental information, etc.), examining the case from multiple perspectives. By comparing and analyzing different text feature sets and paragraph sets, potential evidentiary connections can be discovered. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of the steps of a visual verification method proposed by the present invention; Figure 2A schematic diagram of a module of a visual verification system is provided for the present invention; Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention.
[0018] 610 , processor; 620 , communication interface; 630 , memory; 640 , communication bus. DETAILED DESCRIPTION
[0019] Reference Figures 1 to 3 .
[0020] Example 1 further illustrates a visual verification method proposed by the present invention.
[0021] A visual verification method, comprising the following steps: Collecting the transcript text and associated text and background text associated with the transcript text; Processing the transcript text to obtain a transcript text feature set and an associated paragraph set, segmenting the associated text into paragraphs based on semantic features to obtain an associated text feature set, and segmenting the transcript text into paragraphs based on background text to obtain a background text feature set; Compare and analyze the associated text feature set and the associated paragraph set to obtain the intention association result set; After determining whether the transcript text feature set and the background text feature set are related to each other under the condition of subject relevance, the first intention factor set and the second intention factor set are obtained respectively; After inputting the first intention factor set and the second intention factor set into the intention recognition prediction model, the intention recognition prediction result is obtained. The intention recognition prediction result is combined with the intention association result set to identify the intention of the transcript text and then display the visual evidence content.
[0022] This application first collects the transcript (such as the content of the statement), related text (such as templates for similar cases and legal provisions), and background text (such as information about the case's circumstances and historical handling records). For example, in a dangerous driving case, the transcript includes a description of the incident, related text may include relevant provisions of the Road Traffic Safety Law, and background text may include the time and location information of the body camera video.
[0023] The transcript is processed to extract core semantic features (such as time, location, person, and action verbs). Based on semantic relevance, the transcript is segmented into related paragraph sets (such as the incident description and alcohol test results). Related text is segmented into related text feature sets based on semantic features (such as the sentencing standards paragraph in legal provisions). Background text (such as video recordings) is parsed to extract information such as timestamps and scene descriptions to generate background text feature sets (such as the shooting time and location of body camera videos).
[0024] This approach semantically compares the associated text feature set (e.g., legal clauses) with the associated paragraph set in the transcript (e.g., paragraphs where the parties mention "drunk driving"). This method calculates the frequency of keyword overlap (e.g., "alcohol content" and "driving route") and outputs a set of intent-related results (e.g., determining the strength of the correlation between the parties' statements and the legal clauses). The RoBERTa or BERT model is used to vectorize the text, and cosine similarity is used to calculate semantic match.
[0025] Analyze whether the transcript feature set (e.g., core elements of the case) and the background text feature set (e.g., scenes captured in the video) share the same theme (e.g., "Is the time and location of the crime consistent?"). If the themes are related (e.g., the intersection mentioned in the transcript matches the location of the video), calculate the frequency of semantic keyword overlap (e.g., "Dagui intersection," "alcohol test value 110mg / 100ml"), sentiment (e.g., the degree of cooperation in the parties' statements), and the length of overlapping themes (e.g., the length of the paragraph describing the incident). Combined with historical intent information (e.g., the handling pattern of similar cases), this generates a first set of intent factors.
[0026] If the topics are irrelevant (for example, the time mentioned in the transcript is inconsistent with the video time), the topic distance between the two (such as time difference, location deviation) is calculated and combined with other semantic features to generate the second intention factor set.
[0027] The first intention factor set and the second intention factor set are input into the intention recognition prediction model fine-tuned by the GLM4-9b large model, and the intention recognition results (such as "dangerous driving crime is established" and "the statement is contradictory") are output in combination with the intention association result set.
[0028] Based on the intent recognition results, the system calls up associated visualization resources (such as body camera video clips, evidence photos, and legal text) and presents them in a structured manner according to pre-set rules (such as timeline and evidence chain logic). As the parties recount the incident, the corresponding time point in the body camera video clip is displayed simultaneously, with key information such as alcohol test results and driving routes annotated.
[0029] After determining whether the transcript text feature set and the background text feature set are related to each other under the condition of subject relevance, a first intention factor set and a second intention factor set are obtained respectively, which specifically includes the following steps: When the transcript text feature set and the background text feature set are in a topic-related condition, the transcript text feature set and the background text feature set are compared and analyzed to obtain the first intention factor set; When the transcript text feature set and the background text feature set are in the condition of being irrelevant to the subject, the associated text feature set and the background text feature set are compared and analyzed to obtain the second intention factor set.
[0030] This application first needs to obtain the transcript text feature set, which is a set of key information extracted from the transcript text. For example, the set consists of key semantic information such as the time, place, and behavior of the case extracted from the statement transcripts of the parties involved in the case.
[0031] At the same time, a background text feature set is obtained. The background text may contain environmental information related to the case and other relevant data, from which corresponding feature information is extracted. For example, in a traffic accident case, relevant information such as weather conditions and road monitoring records constitute the background text feature set.
[0032] The transcript feature set and the background text feature set are analyzed to determine whether they are thematically related. For example, if the time and location of the incident described in the transcript matches the time and location of the surveillance footage in the background text, and the subjects involved are consistent, they can be preliminarily determined to be thematically related. Conversely, if there is a clear conflict, such as the transcript stating that the incident occurred during the day, while the background surveillance footage shows it was at night, then they are determined to be thematically unrelated.
[0033] When the transcript feature set and the background text feature set are determined to be thematically related, a comparative analysis is performed on the two feature sets. This comparison process may involve calculating the degree of overlap between the two features, such as counting the frequency of identical keywords and the proportion of similar semantic segments. It is also possible to analyze correlations between features, such as the interaction between character behavior in the transcript and environmental factors in the background text. These comparisons yield a first set of intent factors. This factor set comprises a collection of key factors that, when compared based on thematic relevance between the two text feature sets, can reflect the underlying intent of the text.
[0034] If the subject is not relevant, the associated text feature set is compared with the background text feature set. The associated text feature set may be extracted from other materials related to the case, such as relevant laws and regulations, similar case records, etc.
[0035] For example, the second set of intention factors is obtained by analyzing the differences between the standards stipulated in the associated text and the actual situation of the background text. This factor set reflects the key elements of text intention obtained by combining the associated text and the background text when the transcript text and the background text are not related in topic.
[0036] The transcript text is processed to obtain a transcript text feature set and an associated paragraph set, the associated text is segmented into paragraphs according to semantic features to obtain an associated text feature set, and the transcript text is segmented into paragraphs according to background text to obtain a background text feature set, specifically including the following steps: Extracting semantic distribution features of the transcript text, and segmenting the transcript text into several paragraphs based on the semantic distribution features to obtain a transcript paragraph set; The transcript paragraph set and the background text are segmented into several paragraphs according to their own semantic features to obtain the associated paragraph set and the background paragraph set; After determining the tags of the transcript paragraph set and the associated text according to the semantic importance, the transcript text feature set and the associated text feature set are obtained respectively; The background text feature set is obtained by marking the core intent paragraphs according to their respective semantic importance.
[0037] This application uses natural language processing technology to analyze transcripts from drunk driving cases. For example, from the following sentence: "I dined with friends at XX Restaurant last night and drank about three bottles of beer. Around 9:30 after dinner, I drove back home because I was close to home. I was stopped by traffic police on the way." Semantic distribution features such as "Drinking time (during dinner last night)", "Amount of alcohol consumed (three bottles of beer)", "Driving time (around 9:30 after dinner)", and "Reason for driving (because I was close to home)" are extracted. These features reflect the distribution of key information in the transcript.
[0038] The transcript is divided into different paragraphs based on the extracted semantic distribution features. In the above example, the "Drinking Situation Paragraph" (describing the dining and drinking at the restaurant) and the "Driving Decision and Process Paragraph" (describing the decision to drive after dinner and the process of being caught by the traffic police) can be segmented, thus forming a set of transcript paragraphs.
[0039] The obtained transcript paragraphs are segmented based on semantic features. In the case of drunk driving, the "drinking situation paragraph" can be further segmented into related paragraphs such as "drinking location" (XX restaurant), "amount and type of alcohol consumed" (3 bottles of beer), etc.; the "driving decision and process paragraph" can be segmented into related paragraphs such as "driving time" (around 9:30), "driving reason" (thinking about being close to home), and "circumstances of being checked" (being stopped by traffic police on the road), thus forming a set of related paragraphs.
[0040] Background text may include text recorded by body cameras and surveillance footage from the road where the incident occurred. This background text is segmented based on its semantic characteristics. For example, body camera text can be segmented chronologically into "time segments of traffic police checkpoints" and "time segments of the vehicle involved and interception." Surveillance footage from the road where the incident occurred can be segmented into "time segments of the vehicle involved and interception." Ultimately, this results in a set of background segments.
[0041] Assess the semantic importance of each paragraph in a transcript. In drunk driving transcripts, paragraphs such as "amount and type of alcohol consumed" and "driving time" are crucial for determining whether a driver is driving under the influence and the severity of the offense, and therefore have high semantic importance. By marking these important paragraphs and key semantic information, we generate a transcript feature set, such as explicitly recording key content such as "consumed 3 bottles of beer" and "driving time was around 9:30 PM."
[0042] Related text can include legal documents related to drunk driving, similar case texts, and so on. These related texts are tagged based on their semantic importance. For example, within legal documents related to drunk driving, content such as "blood alcohol content standards for drunk driving" and "corresponding penalties" are semantically important. After being tagged, they form a related text feature set.
[0043] Determine the core intent segments of the background paragraph set. For example, in the background text recorded by the body camera, the segment recording the alcohol test result of the person being stopped is a core intent segment; in the surveillance footage of the road section where the incident occurred, the segment recording the speed and route of the person being stopped is likely a core intent segment. Label these core intent segments to obtain the background text feature set.
[0044] Comparing and analyzing the associated text feature set with the associated paragraph set to obtain the intent association result set includes the following steps: Statistically associate the semantic keyword frequency dataset of each paragraph in the text feature set; Comparing the corresponding semantic keyword coincidence frequencies of the associated paragraph set and the paragraphs in the associated text feature set to obtain a first semantic coincidence frequency dataset; Calculating the proportion of the first semantic coincidence frequency dataset and the semantic keyword frequency dataset to obtain a semantic proportion dataset; When the semantic proportion data set contains a value greater than the preset semantic threshold, the difference exceeding the preset semantic threshold is calculated to obtain the semantic excess data set; After extracting the first intention information of the associated paragraph set, a scene context dataset is set; The intent association result set is obtained by inputting the first intent information, scene context dataset, associated text feature set and semantic excess dataset into the intent association value prediction model.
[0045] In this application, the associated text feature set for drunk driving cases includes legal and regulatory text related to drunk driving (such as the provisions regarding drunk driving in the Road Traffic Safety Law), medical research on the effects of alcohol, and other texts. Semantic keyword frequencies are counted for each paragraph in these associated text feature sets. For example, within a paragraph related to the Road Traffic Safety Law, semantic keywords such as "drunk driving," "alcohol content," "revoked driver's license," and "criminal liability" may be used. Their frequency of occurrence within the paragraph is then counted. For example, suppose "drunk driving" appears five times, "alcohol content" appears three times, and so on, to form a semantic keyword frequency dataset.
[0046] The associated paragraph set is a collection of paragraphs from drunk driving case transcripts, segmented by semantic features, such as "drinking status paragraph" and "driving process paragraph." The semantic keyword overlap frequency of these associated paragraph sets is compared with the paragraphs in the associated text feature set.
[0047] For example, consider the sentence "I drank about three bottles of beer and then drove off." Keywords like "beer" and "driving" in the associated paragraph set are compared with keywords in the Road Traffic Safety Law clauses in the associated text feature set. Assuming there are 10 keywords in the associated text feature set and 5 of them overlap in the associated paragraph set, the overlap frequency is calculated. After comparing multiple paragraphs, the first semantic overlap frequency dataset is generated.
[0048] Calculate the semantic keyword frequency data set from the first step using the first semantic overlap frequency data set from the previous step. For example, if the keyword "drunk driving" appears 5 times in a certain associated text feature set, and the overlap frequency of "drunk driving" in the associated paragraph set is 2 times, then the semantic keyword share of "drunk driving" is 2 ÷ 5 = 0.4. Perform this calculation for all relevant keywords to obtain the semantic share data set.
[0049] Preset a semantic threshold, such as 0.3 (this can be determined based on actual needs and a large number of sample training samples). Check the values in the semantic share dataset. If a value exceeds 0.3, calculate the difference between the values exceeding 0.3. For example, the semantic share of the keyword "drunk driving" is 0.4, exceeding the preset threshold of 0.3. Therefore, the difference between the values exceeding 0.3 is 0.4 - 0.3 = 0.1. Calculate the difference for all keywords exceeding the threshold to obtain the semantic excess dataset.
[0050] Extract primary intent information from the associated paragraphs. For example, in a drunk driving case, the associated paragraphs might mention "I thought drinking this little bit of alcohol was okay," which reflects the driver's lack of awareness of the risks of drunk driving. A scenario context dataset is created based on the scene of the incident, such as the time (night) and location (on the road near the hotel).
[0051] The first intent information, the scene context dataset, the associated text feature set, and the semantic excess dataset are input into an intent association value prediction model (which can be trained using a machine learning algorithm, such as logistic regression or a decision tree) to generate an intent association result set. For example, this can determine the degree of correlation between the party's statement and laws and regulations related to drunk driving, assessing whether the party is aware that their behavior constitutes drunk driving, and other intent association situations.
[0052] When the transcript text feature set and the background text feature set are in a topic-related condition, the transcript text feature set and the background text feature set are compared and analyzed to obtain a first intention factor set, specifically including the following steps: When the transcript text feature set and the background text feature set are under a condition of being related to the subject, the transcript text feature set and the background text feature set are processed to obtain a second semantic coincidence frequency data set and a semantic sentiment tendency value; Counting the overlapping topics of each paragraph in the transcript text feature set and the background text feature set to obtain the overlapping topic dataset; The second intention information of the background text feature set is extracted, and the second semantic coincidence frequency data set, semantic sentiment tendency value, coincidence topic data set, second intention information, scene context data set, intention association result set and transcript text feature set are combined into the first intention factor set.
[0053] In drunk driving cases, the transcript text feature set includes key information about the driver's drinking and driving process, such as "I drank a few bottles of beer and then drove on the road." The background text feature set may come from body camera recordings or surrounding surveillance information, such as "A car with an unusual driving trajectory was discovered on a certain road section at a specific time and was intercepted."
[0054] The two feature sets are compared to identify recurring semantic keywords, such as "beer," "driving," and "traveling." Their co-occurrence frequency in both feature sets is counted to form a second semantic overlap frequency dataset. For example, if the word "driving" is mentioned multiple times in both the transcript and the background text, the percentage of its overlapping occurrences is calculated.
[0055] Determine the emotional attitudes expressed by the parties in the transcript. For example, if a party says, "I didn't realize it was drunk driving at the time; I just thought it was okay since the road was short," natural language processing technology can be used to analyze their emotional tendencies to determine whether they harbored a sense of luck or negligence, and generate a semantic emotional tendency value (which can be quantified into varying degrees of emotional indices, such as -0.3 indicating a slight sense of luck).
[0056] We identify overlapping topics within each paragraph in the transcript feature set and the background text feature set. For example, in a drunk driving case, the description of "driving a vehicle on the road after drinking" in the transcript and the background text recording of "inspecting vehicles on the road after drinking" from the body camera are overlapping topics. We count the corresponding paragraphs in the two feature sets to form an overlapping topic dataset. For example, we record the number of overlapping topic paragraphs and their specific locations.
[0057] Extract information related to the case's intent from background text feature sets (such as body camera recordings). For example, details such as nervousness and avoidance of eye contact during a traffic stop in body camera recordings may indicate that the person is aware that their behavior is illegal, i.e., secondary intent information.
[0058] The second semantic overlap frequency dataset, semantic sentiment tendency value, overlap topic dataset, and second intention information obtained previously are combined with the scene context dataset (including scene information such as the incident occurred at night and the location was on the road near the hotel), the intention association result set (the result obtained through the previous analysis of associated texts and associated paragraphs), and the transcript text feature set to form the first intention factor set.
[0059] When the transcript text feature set and the background text feature set are under a condition of being topic-related, the transcript text feature set and the background text feature set are processed to obtain a second semantic coincidence frequency dataset and a semantic sentiment tendency value, specifically comprising the following steps: When the transcript text feature set and the background text feature set are in a topic-related condition, the second semantic coincidence frequency data set is obtained by counting the frequency of coincidence of corresponding semantic keywords in each paragraph of the transcript text feature set and the background text feature set; Extract the semantic sentiment tendency value of the corresponding paragraph of the second semantic coincidence frequency dataset.
[0060] In drunk driving cases, the transcript feature set includes the parties' descriptions of the events. For example, "I had a party with friends that evening, drank a few glasses of liquor, and then drove home." Semantic keywords such as "evening," "liquor," "driving," and "home" are extracted from this. The background text feature set may come from body camera recordings or intersection surveillance footage. For example, "A car was spotted driving abnormally on a certain road section at night. Upon inspection, the driver showed signs of intoxication." Related semantic keywords such as "night," "driving," and "drinking" are also extracted from this text.
[0061] The semantic keywords of each paragraph in the transcript text feature set and the background text feature set are compared. For example, the transcript text mentions "evening" and the background text mentions "nighttime". The two are semantically similar and are considered to be overlapping; "driving" and "driving" also have semantic relevance and are also considered to be overlapping. The frequency of these semantic keywords appearing together in the two feature sets is counted to calculate the overlap frequency of each keyword. The overlap frequency data sets of multiple keywords are combined to obtain the second semantic overlap frequency data set. For example, the keyword "drinking" appears frequently in both the transcript text and the background text, and its overlap frequency is calculated to be 0.6 (indicating that it appears simultaneously in 60% of the relevant texts).
[0062] Sentiment analysis models can be used to categorize the sentiment within a text into positive, negative, and neutral categories, and assign corresponding quantitative values. For example, if a person's statement reveals a lack of awareness and contempt for the risks of drunk driving, this is judged as a negative sentiment tendency and assigned a sentiment tendency value, such as -0.4 (the range of values can be customized, with negative numbers representing negative sentiment). By performing this analysis on multiple relevant paragraphs, multiple semantic sentiment tendency values are generated, reflecting the emotional attitude information contained within the text.
[0063] When the transcript text feature set and the background text feature set are in a topic-unrelated condition, the associated text feature set and the background text feature set are compared and analyzed to obtain a second intention factor set, specifically including the following steps: When the transcript text feature set and the background text feature set are in a topic-unrelated condition, the topic information dataset is obtained by counting the topics between the corresponding paragraphs in the transcript text feature set and the background text feature set; Among them, the second semantic overlap frequency dataset, semantic sentiment tendency value, overlap topic dataset, second intention information, scene context dataset, transcript text feature set, intention association result set and topic information dataset are combined into the second intention factor set.
[0064] In a drunk driving case, if the transcript feature set mentions that the driver "drank at XX bar and then drove," but the background text feature set (such as surrounding surveillance footage) shows the incident occurred on a highway far from the bar, and the time doesn't match, this indicates that the two topics are unrelated. In this case, it's necessary to count the thematic differences between each paragraph in the transcript feature set and the background text feature set. For example, the transcript may emphasize the drinking scene at the bar and the starting point of the drive from the bar, while the background text focuses on the vehicle's driving status on the highway and the location of the interception. We can determine the themes of these paragraphs and identify information such as "location difference (bar vs. highway)" and "time difference (the transcript time doesn't match the surveillance time)." This information reflecting the thematic differences can be organized into a thematic information dataset.
[0065] The second semantic overlap frequency dataset reflects the overlap frequency of semantic keywords between the transcript feature set and the background text feature set (even if the topics are unrelated, there is still some semantic connection). For example, in a drunk driving case, keywords such as "vehicle" and "driving" may appear in both the transcript and the background text. The overlap frequency of these keywords is counted to form this dataset.
[0066] The sentiment tendency quantified value obtained from the relevant paragraphs of the transcript. For example, if a person complains in the transcript, "I only drank a little bit of alcohol, and the traffic police were too strict," the sentiment analysis model can derive a negative sentiment tendency value, such as -0.3, indicating dissatisfaction and resistance.
[0067] Although the topics of the transcript and the background text are unrelated, there may still be a small number of overlapping topics, such as "road driving". The relevant information of these overlapping topics is counted to form this dataset.
[0068] Information related to the case's intent is extracted from background text feature sets (e.g., body camera recordings). For example, if a person in a body camera recording is evasive when questioned, this may indicate an intention to conceal certain facts, which is secondary intent information.
[0069] This dataset includes scene-related data such as environmental information at the time of the incident, such as weather conditions (whether rain affected driving performance) and road conditions (such as traffic flow on the highway). It also includes key information from the parties' descriptions of the drunk driving incident, such as alcohol consumption and driving routes. This dataset, derived from a previous analysis of the associated text feature set and the associated paragraph set, reflects information such as the degree of intentional connection between texts. This dataset also reflects the thematic differences between the transcript feature set and the background text feature set.
[0070] Example 2 further illustrates a visual verification method proposed by the present invention.
[0071] A visual evidence system, comprising: Collection module: collects the transcript text and the associated text and background text associated with the transcript text; Segmentation module: processes the transcript text to obtain a transcript text feature set and an associated paragraph set, segments the associated text into paragraphs based on semantic features to obtain an associated text feature set, and segments the transcript text into paragraphs based on background text to obtain a background text feature set; The first analysis module compares and analyzes the associated text feature set and the associated paragraph set to obtain the intention association result set; The second analysis module determines whether the transcript text feature set and the background text feature set are related to each other and processes and analyzes them to obtain the first intention factor set and the second intention factor set respectively; Display module: After inputting the first intention factor set and the second intention factor set into the intention recognition prediction model, the intention recognition prediction result is obtained. The intention recognition prediction result is combined with the intention association result set to identify the intention of the transcript text and then the visual evidence content is displayed.
[0072] An electronic device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, a visual verification method is implemented.
[0073] like Figure 3 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute a visual verification method.
[0074] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0075] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform a visual verification method.
[0076] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is configured to execute a visual verification method when the computer program is executed by a processor.
[0077] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0078] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A visual verification method, characterized in that: The method comprises the following steps: Collecting the transcript text and associated text and background text associated with the transcript text; Processing the transcript text to obtain a transcript text feature set and an associated paragraph set, segmenting the associated text into paragraphs based on semantic features to obtain an associated text feature set, and segmenting the transcript text into paragraphs based on background text to obtain a background text feature set; Compare and analyze the associated text feature set and the associated paragraph set to obtain the intention association result set; After determining whether the transcript text feature set and the background text feature set are related to each other under the condition of subject relevance, the first intention factor set and the second intention factor set are obtained respectively; After inputting the first intention factor set and the second intention factor set into the intention recognition prediction model, the intention recognition prediction result is obtained. The intention recognition prediction result is combined with the intention association result set to identify the intention of the transcript text and then display the visual evidence content.
2. A visual verification method according to claim 1, characterized in that: After determining whether the transcript text feature set and the background text feature set are related to each other under the condition of subject relevance, a first intention factor set and a second intention factor set are obtained respectively, which specifically includes the following steps: When the transcript text feature set and the background text feature set are in a topic-related condition, the transcript text feature set and the background text feature set are compared and analyzed to obtain the first intention factor set; When the transcript text feature set and the background text feature set are in the condition of being irrelevant to the subject, the associated text feature set and the background text feature set are compared and analyzed to obtain the second intention factor set.
3. A visual verification method according to claim 2, characterized in that: The transcript text is processed to obtain a transcript text feature set and an associated paragraph set, the associated text is segmented into paragraphs according to semantic features to obtain an associated text feature set, and the transcript text is segmented into paragraphs according to background text to obtain a background text feature set, specifically including the following steps: Extracting semantic distribution features of the transcript text, and segmenting the transcript text into several paragraphs based on the semantic distribution features to obtain a transcript paragraph set; The transcript paragraph set and the background text are segmented into several paragraphs according to their own semantic features to obtain the associated paragraph set and the background paragraph set; After determining the tags of the transcript paragraph set and the associated text according to the semantic importance, the transcript text feature set and the associated text feature set are obtained respectively; The background text feature set is obtained by marking the core intent paragraphs according to their respective semantic importance.
4. A visual evidence method according to claim 3, characterized in that: Comparing and analyzing the associated text feature set with the associated paragraph set to obtain the intent association result set includes the following steps: Statistically associate the semantic keyword frequency dataset of each paragraph in the text feature set; Comparing the corresponding semantic keyword coincidence frequencies of the associated paragraph set and the paragraphs in the associated text feature set to obtain a first semantic coincidence frequency dataset; Calculating the proportion of the first semantic coincidence frequency dataset and the semantic keyword frequency dataset to obtain a semantic proportion dataset; When the semantic proportion data set contains a value greater than the preset semantic threshold, the difference exceeding the preset semantic threshold is calculated to obtain the semantic excess data set; After extracting the first intention information of the associated paragraph set, a scene context dataset is set; The intent association result set is obtained by inputting the first intent information, scene context dataset, associated text feature set and semantic excess dataset into the intent association value prediction model.
5. A visual verification method according to claim 4, characterized in that: When the transcript text feature set and the background text feature set are in a topic-related condition, the transcript text feature set and the background text feature set are compared and analyzed to obtain a first intention factor set, specifically including the following steps: When the transcript text feature set and the background text feature set are under a condition of being related to the subject, the transcript text feature set and the background text feature set are processed to obtain a second semantic coincidence frequency data set and a semantic sentiment tendency value; Counting the overlapping topics of each paragraph in the transcript text feature set and the background text feature set to obtain the overlapping topic dataset; The second intention information of the background text feature set is extracted, and the second semantic overlap frequency data set, semantic sentiment tendency value, overlap topic data set, second intention information, scene context data set, intention association result set and transcript text feature set are combined into a first intention factor set.
6. A visual verification method according to claim 5, characterized in that: When the transcript text feature set and the background text feature set are under a condition of being topic-related, the transcript text feature set and the background text feature set are processed to obtain a second semantic coincidence frequency dataset and a semantic sentiment tendency value, specifically comprising the following steps: When the transcript text feature set and the background text feature set are in a topic-related condition, the second semantic coincidence frequency data set is obtained by counting the frequency of coincidence of corresponding semantic keywords in each paragraph of the transcript text feature set and the background text feature set; Extract the semantic sentiment tendency value of the corresponding paragraph of the second semantic coincidence frequency dataset.
7. A visual verification method according to claim 6, characterized in that: When the transcript text feature set and the background text feature set are in a topic-unrelated condition, the associated text feature set and the background text feature set are compared and analyzed to obtain a second intention factor set, specifically including the following steps: When the transcript text feature set and the background text feature set are in a topic-unrelated condition, the topic information dataset is obtained by counting the topics between the corresponding paragraphs in the transcript text feature set and the background text feature set; Among them, the second semantic overlap frequency dataset, semantic sentiment tendency value, overlap topic dataset, second intention information, scene context dataset, transcript text feature set, intention association result set and topic information dataset are combined into a second intention factor set.
8. A visual verification system, applied to a visual verification method according to any one of claims 1 to 7, characterized in that: include: Collection module: collects the transcript text and the associated text and background text associated with the transcript text; Segmentation module: processes the transcript text to obtain a transcript text feature set and an associated paragraph set, segments the associated text into paragraphs based on semantic features to obtain an associated text feature set, and segments the transcript text into paragraphs based on background text to obtain a background text feature set; The first analysis module compares and analyzes the associated text feature set and the associated paragraph set to obtain the intention association result set; The second analysis module determines whether the transcript text feature set and the background text feature set are related to each other and processes and analyzes them to obtain the first intention factor set and the second intention factor set respectively; Display module: After inputting the first intention factor set and the second intention factor set into the intention recognition prediction model, the intention recognition prediction result is obtained. The intention recognition prediction result is combined with the intention association result set to identify the intention of the transcript text and then the visual evidence content is displayed.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, a visual verification method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method of realizing trial text and trial video synchronization playback and system thereof
CN104869341A
Legal document feature extraction method, related device and storage medium
CN110765889A
Record visualization method and device, electronic equipment and storage medium
CN113468262A
Method for associating multiple evidences in case records based on knowledge graph
CN118761475A
BERT-HiNT and GPT-3-based judgment document generation method, device and system
CN118780252A