GB/T 26237.13-2023 Information technology—Biometric data interchange formats—Part 13:Voice data
GB/T 26237.13-2023 Information technology—Biometric data interchange formats—Part 13:Voice data
Basic Information
Scope
This document specifies a data exchange format that can be used to store, record, and transmit human digital acoustic data (speech) from a single speaker in a single conversation. This data format is used in various speaker identification and verification (SIV) applications, both text-related and text-independent, and makes minimal assumptions about the data collection conditions or environment. It is possible to use the data encapsulated in this format for other purposes, such as automatic speech recognition (ASR), but this is beyond the scope of this document. This document also does not address data processing at the level of voiceprint features or voiceprint models, nor does it address the requirements, devices, or features of specific applications. This document supports non-standard extended data, as well as raw data collected and digitally processed (enhanced) voice data. Any processing of the original input must be described and included in the metadata associated with the speech representation. This document does not address data streaming technology.
This document does not address the timestamping of biometric data stored and transmitted, nor does it address the use of encryption technology to protect the authenticity, integrity, and confidentiality of biometric data.
Information that complies with the format specified in this document can be recorded on machine-readable media and transmitted between systems via data communication.
This document includes provisions for a common voice data exchange format, followed by provisions for an XML schema definition.
This document contains terms commonly used in the fields of speech recognition and speaker identification, as well as terms from other ISO standards.