I have a problem where I generate randomly a dictionary, with a possibly high number of possibilities (say, I have 25'000 possibly different dics). I want to generate an identifier, an ID, for every one of these possibilities. What I want is:
id(x) does not work )My current idea is to use hash functions (although I understand little about it) and do something like this (suppose a dictionary of int/float numbers):
import hashlib
def getID(mydic):
ID = 0
for x in mydic.keys():
# Hash the content
ID = ID + int(hashlib.sha256(str(mydic[x]).encode('utf-8')).hexdigest(), 16)
# Hash the key
ID = ID + int(hashlib.sha256(x.encode('utf-8')).hexdigest(), 16)
return (ID % 10**10)
To my understanding, this should work in most cases, but depending on the actual content of the dictionary and the keys, it's not impossible that two different dics yield the same ID. For example, if I do not hash the keys and two different entries can be "1.0", then I can have a problem.
Do you have anything to suggest, which hopefully does not rely on luck?
Edit: I add a bigger code on what I'm trying to do: it's basically a random parameter optimisation. Code on pastebin
To create an ID, you need to create a non-mutable object. Since keys are unordered, you may need to sort them.
For instance:
mydict = {'a': 1, 'c': 9, 'b': 3}
values = tuple(sorted(mydict.items()))
# -> (('a', 1), ('b', 3), ('c', 9))
Then, you can use your own hash algorithm, for instance with sha256:
import hashlib
def hash_item(m, k, v):
m.update(k.encode('utf-8'))
m.update(str(k).encode('utf-8'))
m = hashlib.sha256()
for k, v in values:
hash_item(m, k, v)
print(m.digest())
# -> b'\xa5\xb42\xee\x03\x07\xbe\x7f\xa2:\xa0\x04a\xf5N\xee4\xba\x9dE%\x1bU\x04V}7\xa8\xda3\x9d\xff'
If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!
Donate Us With