parquet vs csv

GravitySpoiled@lemmy.ml · 2 years ago

parquet vs csv

The Hobbyist@lemmy.zip · 2 years ago

In the deep learning community, I know of someone using parquet for the dataset and annotations. It allows you to select which data you want to retrieve from the dataset and stream only those, and nothing else. It is a rather effective method for that if you have many different annotations for different use cases and want to be able to select only the ones you need for your application.

demesisx@infosec.pub · 2 years ago

How does this differ from graphQL?

ma343@beehaw.org · 2 years ago

Graphql is a protocol for interacting with a remote system, parquet is about having a local file that you can index and retrieve data from in a more efficient way. It’s especially useful when the data has a fairly well defined structure but may be large enough that you can’t or don’t want to bring it all into memory. They’re similar concepts, but different applications

demesisx@infosec.pub · 2 years ago

Thank you!

djnattyp@lemmy.world · 2 years ago

Parquet is a storage format; graphQL is a query language/transmission strategy.